10,000 Matching Annotations
  1. Jul 2026
    1. Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and postmortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

    2. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodents vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

    3. Author response:

      eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

      We thank the Editor and the Reviewers for taking the time to carefully review our manuscript and for providing constructive comments and suggestions, as well as the opportunity to revise our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

      We sincerely thank the reviewer for the positive evaluation of our study and for highlighting the significance of examining synaptic tagging and capture following theta-burst stimulation (TBS) in rodents and non-human primates.

      We agree that TBS is a physiologically relevant induction paradigm and that differences in inhibitory circuit dynamics may also contribute to the species-specific effects observed in our study. As highlighted by Larson and Munkácsy (2015), repeated bursts delivered at theta frequency (~5 Hz) can transiently suppress feed-forward inhibition through GABAB receptor-mediated mechanisms, thereby enhancing postsynaptic depolarization and facilitating NMDA receptor activation. We therefore agree that species differences in inhibitory regulation and burst-evoked depolarization may contribute to the distinct expression of synaptic tagging and capture observed between rats and non-human primates.

      We further agree that analysis of the “depolarization envelope” during sequential bursts may provide additional insight into the relative strengths of early induction mechanisms. We will therefore perform these analyses using the existing recordings and compare the depolarization envelope between rodents and NHPs in the revised manuscript. Following the reviewer’s suggestion, we will expand the Discussion section to acknowledge the potential contribution of inhibitory circuit dynamics and depolarization envelope differences during sequential bursts.

      Importantly, however, we believe that differences in downstream molecular maintenance mechanisms also contribute to these species-specific effects. In support of this, our molecular analyses revealed enhanced recruitment of plasticity-related proteins and transcriptional pathways in NHP hippocampus following TBS, including increased expression of BDNF and PKCζ. These findings suggest that both induction-related network properties and downstream molecular stabilization mechanisms may collectively contribute to the enhanced associative plasticity observed in NHPs.

      We also thank the reviewer for the important point regarding PKMζ antisense experiments and the study by Mei et al. (2011). We agree that the absence of an effect of PKMζ antisense oligodeoxynucleotides does not necessarily exclude a role for persistently elevated PKMζ in the maintenance of theta-burst late-LTP. As demonstrated by Mei et al., BDNF together with theta-burst stimulation can maintain late-LTP in the absence of protein synthesis, potentially through stabilization of PKMζ protein levels by reducing degradation rather than through de novo synthesis. However, these findings are not directly comparable to our study, since our experiments involved theta-burst stimulation alone without exogenous BDNF application. Interestingly, our results suggest species-specific differences in the interaction between BDNF and PKMζ signaling pathways. In rats, TrkB/Fc-mediated blockade of BDNF impaired TBS-LTP maintenance, whereas PKMζ inhibition alone had no significant effect. In contrast, in NHP hippocampal slices, inhibition of either BDNF signaling or PKMζ alone failed to abolish late-LTP, whereas simultaneous inhibition of both pathways disrupted LTP maintenance.

      These findings suggest that endogenous BDNF signaling and PKMζ may operate through partially redundant or compensatory mechanisms, particularly in the primate hippocampus. Therefore, although our findings indicate that de novo PKMζ synthesis may not be strictly required under the present experimental conditions, we cannot fully exclude the possibility that protein synthesis-independent stabilization or maintenance of PKMζ contributes to theta-burst late-LTP maintenance in rodents or NHPs. We will now clarify this point in the revised Discussion section.

      Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      We thank the reviewer for this important and thoughtful comment regarding statistical interpretation and biological replication. We agree that, particularly for electrophysiological experiments where multiple slices may originate from the same animal, the effective sample size for species-level conclusions should be considered at the animal level rather than solely at the slice level.

      In the revised manuscript, we will clearly indicate the number of biological replicates (animals) together with the number of slices contributing to each electrophysiological experiment, as well as the biological replicates used for qPCR and Western blot analyses. We will also clarify whether multiple slices from the same NHP/rat contributed to the same experimental condition. These details will be incorporated into the figures and figure legends wherever appropriate.

      In addition, we will perform animal-level analyses by averaging slice responses within each animal prior to statistical comparison and, where appropriate, apply hierarchical or mixed-effects statistical models to account for the nested structure of slices within animals.

      We acknowledge that the number of non-human primates (NHPs) available for this study was inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with primate electrophysiology and tissue collection. Consequently, achieving sample sizes comparable to rodent studies is often not feasible in NHP research. Nevertheless, to further strengthen the biological robustness of the findings, we are currently in the process of obtaining additional NHP brain samples and plan to repeat key experiments in an additional 3-4 animals. We believe these revisions and additional experiments will substantially strengthen the statistical rigor and overall interpretation of the study.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      We thank the reviewer for this insightful comment regarding the mechanistic interpretation of the STC findings. In the present study, we selected the 30 min interval based on well-established classical STC paradigms in rodents, where this interval reliably falls within the effective tagging and capture window. Using this experimentally validated interval allowed us to directly compare whether TBS is sufficient to support STC in primates versus rats under equivalent experimental conditions. Accordingly, the primary objective of this study was to determine whether TBS-induced STC varies across species, rather than to comprehensively define the temporal dynamics of the tagging window.

      We agree, however, that the current experiments do not distinguish whether the primate-specific effect reflects prolonged tag persistence, enhanced plasticity-related protein (PRP) synthesis, altered capture efficiency, or a shifted temporal window. Addressing these possibilities would indeed require systematic temporal interval analyses (e.g., ±15, ±30, ±60, and ±90 min), which represent important future directions. Such experiments are particularly challenging in non-human primates because the availability of primate tissue and experimental resources for large-scale electrophysiological studies remains limited and is currently beyond our experimental capacity due to substantial ethical, logistical, financial, and technical constraints.

      Nevertheless, we fully agree with the reviewer that these experiments are important for advancing the mechanistic interpretation of the findings. Similar temporal analyses have recently proven informative in our rodent studies (Chong YS, Ang SR, Sajikumar S. Commun Biol. 2025;8:553). Importantly, we are currently in the process of obtaining additional non-human primate samples and plan to extend the present work by examining an additional 60 min temporal interval to further characterize the temporal properties of synaptic tagging and capture in non-human primates.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and post-mortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      We thank the reviewer for raising this important and insightful point. We agree that differences in developmental stage between the experimental groups represent an important consideration when interpreting potential species-dependent effects. In the present study, rat experiments were performed in 5-7 week-old animals, whereas non-human primate (NHP) tissues were obtained from 5-7-year-old monkeys. This difference largely reflects the practical, ethical, and logistical constraints associated with NHP research and tissue availability. We acknowledge that these ages are not developmentally equivalent and that maturation state may influence BDNF signaling, protein synthesis capacity, synaptic plasticity thresholds, and transcriptional responses relevant to late-LTP and STC mechanisms.

      We also recognize that differences in euthanasia procedures, tissue extraction, slice preparation, and postmortem handling between rodent and primate tissues may influence tissue physiology and electrophysiological properties. Although extensive care was taken to optimize tissue viability and maintain stable recordings within each species, these variables cannot be completely excluded as contributing factors to the observed differences.

      Accordingly, we will revise the Discussion section to more explicitly acknowledge these limitations and clarify that our findings support potential species-dependent differences under the present experimental conditions, rather than definitive intrinsic species-specific mechanisms. Nevertheless, despite the inherent challenges associated with NHP electrophysiological studies, we believe that the present findings provide an important initial framework for understanding the translational relevance of synaptic tagging and capture mechanisms across species.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      We thank the reviewer for this important comment regarding statistical reporting and interpretation. We agree that the repeated occurrence of identical exact p-values in several nonparametric analyses reflects the relatively small sample sizes and the discrete nature of the statistical distributions. This issue is particularly relevant for the NHP experiments, where biological replication is inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with obtaining and processing primate tissue.

      In the revised manuscript, we will provide exact n values for all comparisons, including the number of biological replicates (animals) and slices where applicable. We will also include additional statistical details, including effect sizes and confidence intervals where appropriate, to improve transparency and facilitate interpretation of the reported findings. Furthermore, we are currently in the process of obtaining additional NHP samples and will attempt to include more biological replicates in the revised version to further strengthen the robustness of the analyses.

      We also agree that the issue of multiple testing should be addressed more explicitly, particularly because multiple genes and proteins were examined. In the revised manuscript, we will clearly state the statistical correction methods applied for multiple comparisons where appropriate. For analyses in which corrections were not applied, we will provide justification, noting that several experiments were based on hypothesis-driven candidate targets rather than exploratory large-scale screening analyses. These statistical considerations will be clarified in the Methods and Results sections.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

      We thank the reviewer for this thoughtful and balanced assessment of our work. We agree that the present data primarily support the conclusion that, under the specific experimental conditions examined, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than that observed in rat slices. We also agree that broader interpretations regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and potentially redundant BDNF/PKMζ-related mechanisms require additional mechanistic investigation and experimental validation.

      Accordingly, we will moderate these interpretations throughout the revised manuscript and clearly state that these conclusions remain preliminary. We will further emphasize that additional experiments, including increased biological replication, expanded temporal analyses, and further mechanistic investigations, will be necessary to more conclusively define the basis of the observed species-dependent differences. Within our current experimental capacity, we are actively working to obtain additional non-human primate samples and plan to incorporate additional biological replicates and key follow-up experiments in the revised version to further strengthen the robustness of the findings.

      At the same time, we believe the present study provides an important initial contribution to an understudied area by directly examining synaptic tagging and capture mechanisms in the primate hippocampus. Given the limited availability of non-human primate electrophysiological data in the field, these findings may offer a valuable framework for future studies investigating the translational and evolutionary relevance of associative synaptic plasticity mechanisms across species.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodent’s vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

      We thank the reviewer for this thoughtful and insightful comment, as well as for the encouraging appreciation of our long-duration plasticity recordings and associative plasticity experiments, which are both technically demanding and time-intensive. We fully agree that interpretation of cross-species differences in synaptic plasticity requires careful consideration of multiple biological and environmental variables, including circadian state, enrichment conditions, strain differences, diet, lighting conditions, and species-specific behavioral ecology.

      Regarding the specific concern related to circadian phase and sleep-wake state, the reviewer raises an important point. Rats are nocturnal animals, whereas macaques are diurnal, and hippocampal plasticity mechanisms are known to be influenced by circadian rhythms and sleep-dependent regulation of synaptic proteins and signaling pathways. Previous studies have demonstrated modulation of LTP, synaptic tagging and capture and protein synthesis in rats across normal sleep-wake cycles. We therefore agree that these factors may influence plasticity outcomes and should be carefully considered in comparative studies.

      Studies have further shown that theta frequency is highly sensitive to sleep-related manipulations. Specifically, theta frequency decreases immediately after sleep, remains elevated during sleep deprivation, and rapidly declines following recovery sleep. In aged animals, these effects appear comparatively attenuated, suggesting reduced sleep-dependent modulation of theta dynamics with aging. Therefore, disruption of normal circadian or sleep-wake patterns may significantly alter theta activity and associated plasticity mechanisms within a species and may not accurately reflect physiological baseline states (Utku Kaya et al., 2026).

      In our experiments, recordings from rats and macaques were performed during their respective active phases under standardized laboratory housing conditions, and we will further clarify these details in the revised Methods section. Nevertheless, we acknowledge that circadian state and related physiological variables cannot be completely excluded as contributing factors to the observed differences between species.

      More broadly, we agree with the reviewer that the present study does not permit definitive conclusions regarding universal “rodent versus primate” rules of synaptic plasticity. Our intention was not to propose a generalized dichotomy between rodents and primates, but rather to report that, under the experimental conditions used here, SC-CA1 TBS-LTP and associated synaptic tagging mechanisms differed between rats and macaques. We agree that broader evolutionary or cognitive interpretations would require systematic comparative analyses across multiple species, including both nocturnal and diurnal rodents as well as diverse primate species. Such studies would provide a stronger framework for distinguishing conserved versus species-specific mechanisms of plasticity.

      At the same time, we believe the present findings remain important because they provide one of the first direct experimental comparisons of SC-CA1 TBS-LTP-associated plasticity mechanisms between rodents and non-human primates under controlled ex vivo conditions. Although the interpretation should be done cautiously, the observed differences raise the possibility that certain metaplastic or protein synthesis-dependent mechanisms may not be fully conserved across species. Accordingly, we will revise the Discussion section to better emphasize the exploratory and comparative nature of the study, while explicitly acknowledging the limitations and potential confounding factors highlighted by the reviewer.

    1. eLife Assessment

      This important study assessed the replicability of a selection of lab-based biomedical experiments in papers published by authors based in Brazil. The study adds a unique perspective to the literature on replication, and provides rich data on the approach taken, the outcomes, and the challenges involved in conducting large-scale crowd-sourced research. The evidence supporting the claims is convincing, but there is scope for clarifying the presentation of the results and extending the discussion section.

    2. Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      (2) The article appears to oscillate between:

      i) a description of the approach and the inherent challenges of such a multicenter replication program.

      ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections. Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers

    3. Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was under represented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX). In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this. If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications. I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean. An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made. In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

    4. Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      We thank the reviewer for the comments.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      We can try to simplify the description of protocols at specific points for example, by providing an overarching description of the study design in the beginning of the Methods, rather than citing our previous eLife paper (Amaral et al., 2019), as suggested below. The methods are indeed quite extensive, but the this may be inevitable in a large-scale project such as this and we note that Reviewer #2 thought that part of the supplementary material should be incorporated back in the main text, which is a suggestion in the opposite direction. It may thus be hard to strike a balance between readability and comprehensibility that can address both reviewers’ opinions.

      (2) The article appears to oscillate between:

      (i) a description of the approach and the inherent challenges of such a multicenter replication program

      (ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      There is a bit of redundancy between tables and text, but this was intentional to make both of them self-explanatory. We also think stating the results in the text can allow us to make each of the replication criteria clearer, a concern that was also mentioned by the reviewer.

      As for requiring particular expertise in statistics for understanding, we mostly disagree. The main results (Tables 1 and 2, Figure 2) are expressed as percentages, and the only statistical concepts needed for interpreting these results are understanding prediction and confidence intervals. For this, we could provide a bit more guidance on their interpretation in the Methods section. Beyond that, most of the secondary results (e.g. Figure 3 and Figure 4) involve linear correlations, which is about as simple as statistical analysis gets.

      Of the results presented in the main manuscript, only Table 3 contains anything beyond percentages and correlations. We do agree that the meaning of each ratio in this table could be more clearly described, but there are essentially no expert-level statistics involved in their calculations.

      Other than that, the main statistical issues are the ideal way to aggregate the results from different replications for which we use different strategies for robustness purposes. However, all of these results are already in the supplementary material, so we don’t feel they interfere to much with the readability of the main manuscript.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections.

      This is indeed a good idea, and we plan to include an initial overarching description of the project in the Methods section of the revised manuscript.

      Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      Again, this is the opposite of what was suggested by Reviewer #2, so we would rather keep the Methods section more or less at its current level of detail.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      The reviewer is correct in his interpretation. Evaluating the main findings of articles or cleaning a field of wrong statements was never a goal of our study (and we were clear about this from the start). Our aim with the project was metascientific (i.e. evaluate the reproducibility of biomedical experiments with a set of common methods) rather than driven by a particular interest in the findings themselves. This is reflected by our choice of selecting experiments from a random sample of articles from multiple fields, rather than filtering by area of interest or importance. It also underlies our choice to evaluate experiments rather than claims, as this was more statistically tractable and potentially more objective as a meta-research goal.

      To be clear, we don’t feel this approach is inherently better or worse than evaluating claims in the literature, as in the Drosophila immunity article case (i.e. Westlake et al., 2026), which is also an important goal. They are merely approaches that answer different questions. Ultimately, we probably made our choice based on (a) our expertise/interest in meta-research rather than in the fields the replications stemmed from and (b) an attempt to engage Brazilian researchers in the project in a way that was non-confrontational and minimized backlash from their peers. We feel this was valuable for many of the lessons learned, although it also meant learning less about the research findings in question.

      Even though this was not a goal of the study, there is some knowledge obtained about the findings that is indeed largely absent from the current manuscript. We do not feel the current format allows for much discussion of 45 different findings, but we do have plans to address these in future articles (as outlined in our response to point 5). In the meantime, qualitative descriptions of each experiment can be found at https://osf.io/w5z9a. This is already mentioned in the Methods but could be reiterated in the results as well.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      We did not go too deep into that finding because we are publishing a separate article focused on the prediction project, which should look into factors that correlate with prediction accuracy, both at the level of predictors (e.g. research field, career level) and of individual predictions (e.g. information taken into account for each answer). We also feel that, given the multiplicity of predictors in the prediction analyses, these findings are a bit tentative, as the strongest predictors may be subject to effect size inflation from the “winner’s curse” effect (as outlined by Reviewer #2). We can try to emphasize it a little more in the discussion (although it already merits a whole paragraph on pages 23-24), but we feel we would be able to discuss it more critically in a follow-up article.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers.

      Indeed, some of our results did not fit this overarching analysis and were left for future publications. One of them is already available as a preprint, while the others are currently in preparation. Specifically, other results from the project should be spread about across five different articles.

      (a) A narrative article focused on challenges and lessons learned with the project, already published as a preprint at https://osf.io/preprints/metaarxiv/8y3tg_v1 (Amaral et al., 2026).

      (b) An article analyzing the prediction survey and markets results in detail (following the pre-analysis plan detailed in https://osf.io/6av7k/files/pjhgd and adding some exploratory analyses on prediction rationales).

      (c) Three articles describing the results of specific experiments with each experimental method (MTT, PCR, elevated plus maze) along with a discussion of aspects inherent to the method that seem to influence reproducibility.

      We can add this information more explicitly to the Methods section, including the links to the papers that have already been published at the time the manuscript is revised.

      Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was underrepresented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Thanks!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      We once more thank the reviewer for the compliments.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      We do acknowledge that the article currently includes a lot of supplementary material. This includes both supplementary figures/tables relating to the paper and many supplementary methods files (mostly hosted at the Open Science Framework). However, we also note that this is already a rather long paper as it stands and that Reviewer #1 has made the opposite suggestion of simplifying it. Thus, it may be hard to strike a balance that will suit all preferences, and we feel that maybe our attempt has landed somewhere in the middle of both reviewers’ ideal versions of the paper.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      We can try to engage with some of the above-mentioned literature in more depth in particular replication studies from other fields (some of which have appeared after our preprint (e.g. Tyner et al., 2026) and with the risk of bias and transparency literature (e.g. Serghiou et al., 2021). That said, we note once more that the article (and the Discussion section) are already quite long, and that analyzing each of these articles in depth is likely to be unfeasible.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Our original plan was to use p values as a predictor (see protocol at https://osf.io/9rnuj), but we later realized this was inadequate as it did not account for effect direction (i.e. significant effects in the opposite direction as the original may yield low p values, but this should not count as replication success). We thus switched to t values to be able to assign positive and negative signs depending on effect size direction. We note that, as we are using non-parametric Spearman coefficients (in which the module of t correlates negatively with the p value), the two approaches are effectively equivalent when original and replication effects have the same direction. This change was accounted for and justified in our list of protocol deviations at https://osf.io/9hj7t.

      Effect size (in relative terms) is already being used in the second predictor in the analysis (i.e. effect size decrease), as our idea was to use one significance-based predictor and one effect size-based predictor, to match what was done for the replication rates). We feel that using relative effects (e.g. response ratios) by themselves may not be as adequate, as for experimental methods with large coefficients of variation and/or low sample sizes (especially PCR ones), one can find large relative effects that are nevertheless far from statistical significance. This also makes relative effects not very commensurable between methods.

      We do believe there is a fair argument, however, to use standardized effect sizes as an alternative to t values (i.e. difference measured in standard errors of the mean) to measure significance/evidence strength. As some replications ended up underpowered, low t values may sometimes be due to insufficient statistical power/low sample size rather than replication failures. Using standardized effect sizes is not devoid of pitfalls (e.g. they can be quite variable when sample size is low), but it is worth doing as a robustness analysis.

      That said, there are a few statistical issues to be decided on how to calculate this (e.g. whether studies should be meta-analyzed using standardized mean differences rather than relative ones for this purpose, or whether an analog of the standardized effect size should be calculated for the log ratio of means). We would have to look more carefully into the multiple possibilities to decide on the best approach (and we do accept suggestions!).

      In the meantime, we note that running the prediction analysis using only experiments with ≥80% power yields a slightly higher correlation of t scores with researcher predictions (ρ = 0.49, p = 0.005), so we do not think that these underpowered experiments affect the trend too much. If anything, they could be masking a higher correlation between researcher predictions and replicability.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      We agree that we should make the description more precise (e.g. “reaching the same results when analyzing a set of data in the same way” for reproducibility and “finding similar results with new data collected under similar conditions” for replicability). We will update these definitions in the revised manuscript.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      We would argue that the reader would be correct in this case: the argument is a bit speculative. It does go in the direction of what is generally accepted within the field (i.e. that publication pressure can lead to lower reproducibility for a range of factors), but we’re not sure this connection has been demonstrated empirically, except for indirect evidence (such as the lower reproducibility in papers stemming from top institutions and “trophy journals” in, the higher frequency of positive results in US states with more researchers in Fanelli, 2010, or the higher number of problematic images for highly productive researchers in some countries in Fanelli et al., 2022. We could cite this evidence in the introduction and make the speculated connection more explicit, perhaps adding modeling work as well (e.g. Ioannidis, 2005; Smaldino & McElreath, 2016) to explain why this could be the case. But essentially, our opinion is that the connection remains a speculation.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      We can offer a more detailed description of the recruitment process (e.g. number and distribution of lectures, social media strategy used, etc.), although we would rather do this in a supplementary document so as not to make the Methods section even lengthier. We note, however, that we never aimed to recruit a “representative sample” of labs from the country: we were busy enough trying to get enough labs for the project to happen, and aware that the call would be inevitably biased by our own communication capabilities and personal networks.

      That said, the response rates for different regions of Brazil do generally match the distribution of research labs and graduate programs within the country (with some distortions likely caused by our personal networks, such as the large number of labs in Rio de Janeiro state), and seem to indicate a rather wide dissemination of the call. One way to visualize this would be to present the distribution of corresponding articles from the original studies selected for the replication (or even from the whole sample of articles obtained for experimental selection) along with the distribution of labs at different stages of the project in Figure S3, which generally show similar patterns. This would actually lend support to our statement that “the population of labs that performed replications was largely similar to the one that produced the original results” in the discussion.

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      We tried to find the combination of methods that would maximize the number of labs that would be included in the project. This is explicitly stated in our Methods Selection document at https://osf.io/qxdjt, but could be stated more explicitly in the paper as well.

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      We initially used single screening by three different reviewers (see https://osf.io/6av7k/files/u5zdq for criteria), as we were merely looking for a sample of experiments; thus, comprehensive inclusion of all eligible studies was not a priority. After this initial screening step, inclusions were confirmed in a consensus meeting with the three reviewers involved.

      Data extraction was also done by a single individual, but the resulting data led to a protocol that was later checked by two reviewers who had access to the paper and were explicitly oriented to judge whether the protocol consisted in a valid replication. Thus, discrepancies between what was in the paper and what was included in the protocol could potentially be flagged at these stages (as they were in many cases). We do note, however, that this is likely not as effective to prevent errors as having data extracted independently, as reviewers may overlook mistakes more easily when comparing two documents rather than extracting data anew. We did find that some errors in extraction slipped by, such as an MTT experiment where treatment concentration was inadvertently changed from mM to μM in a particular protocol step; this was picked up and corrected by 2 out of the 3 labs, but not by the third one, leading the latter replication to be invalidated.

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      As stated in the manuscript, we initially capped experiments at a predicted cost of R$ 5.000 (around USD 1336 at that time), considering reagent cost alone (as equipment and labor was provided by labs), as mentioned in the manuscript. Exclusion rates for that reason were 12/74 (16%) for MTT experiments, 36/132 (27%) for PCR ones and 4/40 (10%) for EPM ones. This is stated at

      This turned out to be an underestimation in many cases, especially as it did not account for pilot experiments, need for repetition, etc; thus, many experiments ended up costing considerably more than that ceiling. As we had included a contingency fund for those cases which we expected would occur , we avoided removing experiments from the sample for this reason as much as possible. Nevertheless, one elevated plus maze experiment ended up not being replicated for cost reasons, as the necessary rat strain was provided by a single facility in the country, meaning that a large number of rats would have to be acquired and transported to all labs at a cost that we were not able to cover.

      As these costs were covered by the coordinating team, we do not feel that this is likely to underlie the reduction in geographical coverage. Other reasons related to lab structure could have led to labs in less well-resourced regions to leave the project, but they probably has nothing to do with the experiments selected.

      That said, the cost cap does mean that the selection of experiments is not completely representative of the literature, but is enriched in relatively cheap and simple experiments which were able to perform (which was our next step for selecting the final sample of experiments. Exclusion rates due to lack of lab expertise and/or infrastructure to perform the experiment were 21/56 (37%) for MTT experiments, 67/89 (75%) for PCR ones and 7/34 (21%) for EPM experiments.

      We will try adding some of this information to the flowchart in Figure 1, as we agree it provides more context on the representativeness of the selected experiments.

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      The scale ranged from 1 (No relevant differences) to 5 (Very relevant differences that prevent considering the study as a direct replication). This scale was used for both the lab and the validation committee scores, and is described at https://osf.io/xgth2 (debriefing protocol) and https://osf.io/e3fjg (validation protocol).

      For the validation committee, we did use a threshold (any score of 4 or a sum of scores of 10 or more among 3 evaluators) to decide what had to be discussed to decide on inclusion, as mentioned on Page 7 of the Methods. For the labs, we used no threshold labs answered the protocol deviation question as a scale, but the decision of whether to consider the study a valid replication or not was not tied to this score.

      We can make both of these points (meaning of the scale and connection to lab’s decision to consider the replication valid) clearer in the Methods section.

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      For the initial analysis of justifications, one reviewer read all answers and flagged those that seemed to concern reproducibility of the methods (e.g. “we replicated the protocol exactly as planned”) rather than results reproducibility (e.g. “effects went in the opposite direction”). We then revised these answers among the whole coordinating team to decide whether we should contact the lab asking them to revise them. We can add this information to the Methods section.

      For classifications of the justification into categories (i.e. Table S7), justifications were classified by two independent reviewers based on categories created after an initial inspection of the data, and discrepancies were resolved by consensus. We can add this information to the table legend.

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      After we extracted data from the lab spreadsheets and summarized the results by code, labs received the results by e-mail and were asked to fill in a form on whether the results were in agreement with what they had found (see details at https://osf.io/nfr6y). Discrepancies in results at least 1 experiment were noted by 36% of the 53 (out of 56) labs that responded. Many of these stemmed from the coordinating team misunderstanding issues such as group identity or experimental unit identification in the spreadsheet. Others had to do with different ways to perform calculations (e.g. relative gene expression or % time spent in open arms). In some cases, simple errors in data transcription or typos caused the discrepancy.

      We were also surprised (and concerned) by the number of experiments in which we later found data errors that were not detected by this process (e.g. 18% of total). Our best understanding of this is that not every lab checked the results with the necessary care, as some errors were quite obvious, as in experiments in which sample size was different, or in which group labels were reversed. Ultimately, agreeing with a form that says “did you find any discrepancies?” may have been performed as a box-ticking exercise with little attention, and was probably not the ideal way to check data which led us to start reviewing results in live meetings afterwards. This is discussed in more detail in our challenges article (Amaral et al., 2026)

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX).

      R 4.5.1 was used for the analysis. We can add this information (which was present in the data repository in the R session info.txt file) and provide the R reference in the manuscript as well.

      In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this.

      If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Yes, we did use the escalc() function for this calculation (for both the replications and the original effect sizes). We can mention this in the manuscript.

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications.

      We calculated the coefficients of variation as the pooled SD divided by the mean of both group means. The reviewer is correct about the possibility of small-sample effects in this case (which we were not aware of). We will thus look into the possibility of implementing this via the escalc () function in the analysis of the revised manuscript.

      We also acknowledge that this could be a source of bias in the comparisons between original and replication CVs (albeit likely a minor one). That said, we note that sample sizes are not always larger in the replication for some experiments with large original effects, power calculations sometimes yielded lower sample sizes in the individual replication, albeit infrequently. On average, though, replication sample sizes were indeed larger.

      I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean.

      This is indeed the case; that said, the CV of the original effect also has random error relative to the true population CV and in that case, there is no way to estimate the uncertainty, as we have a single measure of that parameter. So there is probably no way around ignoring uncertainty in this case.

      We also note that we are looking for evidence of systematic CV inflation across all experiments (rather than for a statistically robust comparison between the CVs of any individual replication). For the sake of measuring this systematic inflation, the use of multiple experiments does allow us to estimate variability at the experiment level which should incorporate the lower-level variability between individual replications if this is not included in the model. Thus, we do not feel that our procedure introduced a systematic bias in the analysis at the experiment-level (although one could argue that it may lead to less precision).

      An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      We thank the reviewer for this suggestion, which indeed seems like an option in this case. We will look into this possibility, although we cannot guarantee at the moment that we will implement it, as we were not previously familiar with the method and will have to study it in more detail.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Yes, but there are likely different reasons for that. Labs leaving after being included may have been due to those in less privileged regions of Brazil (e.g. the northern and western regions of Brazil, generally speaking) having more difficulty in persisting in the project. That said, most of the “disappearance” happens between registration and inclusion which usually has to do with the labs not working with the methods that were ultimately included in the project. We also note that most of the states that lose representation were those that had a single lab to begin with, which may make the visual pattern more striking than the actual trend (as states in the South/Southeast also lose labs, but don’t disappear from the map).

      We note again that we never planned to achieve geographical representativeness when recruiting the labs on the contrary, we were aiming to maximize the number of available labs to run the project. That said, we do agree that for the sake of examining whether the population of labs is similar to the one that generated the original experiments (a claim that we do make in the discussion), this representativeness is important to assess. Once more, to allow the reader to evaluate this, we plan to add an additional map to Figure S3 to describe the Brazilian states where the original experiments came from (based on corresponding author affiliations) in which a similar bias towards the South and Southeast Region can be observed.

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      We agree that this would be useful information, and can experiment with the possibility, but our feeling is that the figure will likely become too noisy in cases where the 95% CIs overlap (which are quite frequent). If this is indeed the case, an option to allow the reader to examine this would be better to add an explicit link to the forest plots for each individual experiment (https://osf.io/sx9gv) in the figure legend.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      We selected the highest correlation values for each continuous outcome (t score and lnRR) and presented these separately in the text. This is a systematic way to perform the selection, but is obviously subject to the “winner’s curse” effect. We agree that adding both metrics for each predictor would be a fair way to keep this in perspective for the reader, but we would have to think about how to do this without sounding too confusing (as results for the two main outcomes are quite different).

      We do note, however, that the outcomes are indeed different and are expected to vary independently in some cases. For the correlation with replication probability predictions, for example, the effects in opposite directions would likely be expected, as larger original effect sizes will likely lead to larger probabilities to be assigned, but also to a higher possibility of effect size decrease. This low correlation between outcomes is probably something that should be pointed out and discussed in the revised manuscript.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made.

      We had given this some thought when writing the manuscript – but ultimately opted not to include confidence intervals for our replication percentages and to use the replication rates as descriptive measures only (as done in other replication studies such as (Errington et al., 2021).

      Even though we aimed for our sample of original experiments to be as systematic as possible, it is ultimately constrained by many factors (the choice of methods, the particular expertise of the labs, etc.) thus, adding confidence intervals represents the uncertainty around the replication rate of a very specific population of experiments, which is not directly comparable to those included in other replication efforts in any case.

      We will reconsider whether we should include confidence intervals for replication rates: although doing this for every replication rate in Table 1 and Table 2 may end up being too much information, it could probably be done at least for the replication rates of the main analysis in the text. We note that calculating confidence intervals for percentages is straightforward, requiring only the numbers that are in the table thus, any reader that wants to estimate uncertainty for those rates should be able to do it easily.

      We will also point out the uncertainty around the percentages mentioned in the discussion when comparing our replication rates with those of other studies, which we agree is an important issue to touch on.

      In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Both correlations in that case are non-parametric (e.g. Spearman’s ρ), so they cannot be directly transformed into Pearson’s r without making assumptions about the distribution (which we would probably avoid doing given the very marked outlier in our own). We can calculate a non-parametric confidence interval for our own correlation coefficient by resampling, but we will have to investigate whether this can be done using the available data from (Errington et al., 2021) (which is probably the case if effect sizes for all experiments have been shared).

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      It is indeed interesting, and seems to confirm an intuition that has long been present in the reproducibility field, but actually has little evidence to support it: if anything, there is evidence in the opposite direction in psychology (Youyou et al., 2023), although they looked at cumulative publication number, while we used number of publications in a fixed interval.

      We can expand a bit further on that finding: that said, we do note that the correlation is relatively weak and has a p value of 0.04. Thus, given the multiplicity of predictors would not be that unlikely to occur by chance, even though it seems intuitive. Thus, even though the relationship seems intuitive, we think it should be considered tentative at best and would refrain from discussing it in too much detail.

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Yes, that is what we mean by “incorrect labeling of error bars” (as can be grasped from the cited references).

      We can perform this regression, which seems relatively straightforward to do. That said, we note that another likely cause for outliers at least for cell line studies would be the use of different (and eventually inadequate) experimental units (e.g. having error bars that represent technical replicates of the same measurement rather than truly independent experiments). We suspect that this may have an even greater effect in terms of causing error bars not to express the same thing and the regression will not help in differentiating the two causes.

      We should also note that different types of experiments may be expected to have very different SDs, so the regression is likely to have a lot of error associated with it. In particular, it’s probably worth doing separate regressions for each method, to account for the likely difference in CVs between animal and cell line experiments, for example. This could also help tease apart the two causes above, as the experimental unit problem mentioned above will likely only be observed for cell experiments.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

      Thanks!

      Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      We thank the reviewer for the compliments. Again, a more extensive list of insights can be found in our challenges article (Amaral et al., 2026), which we will cite in the revised version.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

      We acknowledge the missing replications as a weakness, and we hope we have made that point clear in the discussion.

      Concerning the Brazilian research ecosystem, we could try to explore this in more detail in the introduction. In particular, we believe that a better understanding of the Brazilian academic system, including its regional disparities and the general composition of its workforce (which is largely composed of undergraduate and graduate students), can be useful in interpreting some of the findings.

      We can try to provide a bit more context at the end of the introduction (perhaps between the last 2 paragraphs, which would also address a point made by Reviewer #1), and also in different points of the discussion including those comparing replication rates with other studies or discussing infrastructural difficulties, some of which may be specific to the Brazilian context (such as difficulties in acquiring specific reagents or licenses). Still, we reiterate that, due to the lack of studies with comparable samples in other regions, we cannot tease apart the factors that are specific to Brazil from those affecting lab biology as a whole from the data alone.

      References:

      Amaral OB, Neves K, Wasilewska-Sampaio AP, Carneiro CF. 2019. The Brazilian Reproducibility Initiative. eLife 8:e41602. DOI: https://doi.org/10.7554/eLife.41602

      Amaral OB, Valério B, Carneiro CFD, Mota GPS, Neves K, Abreu M, Tan PB. 2026. Challenges for building up confirmatory science in lab biology: lessons learned from the Brazilian Reproducibility Initiative. MetaArXiv, DOI: https://doi.org/10.31222/osf.io/8y3tg_v1

      Errington TM, Mathur M, Soderberg CK, Denis A, Perfito N, Iorns E, Nosek BA. 2021. Investigating the replicability of preclinical cancer biology. eLife 10:e71601. DOI: https://doi.org/10.7554/eLife.71601

      Fanelli D. 2010. Do pressures to publish increase scientists’ bias? An empirical support from US states data. PLoS One 5:e10271. DOI: https://doi.org/10.1371/journal.pone.0010271

      Fanelli D, Schleicher M, Fang FC, Casadevall A, Bik EM. 2022. Do individual and institutional predictors of misconduct vary by country? Results of a matched-control analysis of problematic image duplications. PLoS One 17:e0255334. DOI: https://doi.org/10.1371/journal.pone.0255334

      Ioannidis jpa. 2005. why Most Published Research Findings Are False. PLoS Medicine 2. DOI: https://doi.org/10.1371/journal.pmed.0020124

      Serghiou S, Contopoulos-Ioannidis DG, Boyack KW, Riedel N, Wallach JD, Ioannidis JPA. 2021. Assessment of transparency indicators across the biomedical literature: How open is open? PLOS Biology 19:e3001107. DOI: https://doi.org/10.1371/journal.pbio.3001107

      Smaldino PE, McElreath R. 2016. The natural selection of bad science. R Soc Open Sci 3:160384. DOI: https://doi.org/10.1098/rsos.160384, PMID: 27703703

      Tyner AH, Abatayo AL, Daley M, Field S, Fox N, Haber NA, Hahn KM, Struhl MK, Mawhinney B, Miske O, Silverstein P, Soderberg CK, Stankov T, Abbasi A, Aberson CL, Aczel B, Adamkovič M, Albayrak N, Allen PJ, Andreychik M, Awtrey E, Axxe E, Azevedo F, Bader MD, Bago B, Bailey J, Bakker M, Banik G, Banks GC, Baskin E, Batruch A, Beatteay A, Behr SM, Berente N, Berry Z, Białkowski J, Bodroža B, Boeschoten L, Bognar M, Bokhove C, Bonfiglio D, Bouwman R, Brady TF, Braithwaite SR, Briceño Jiménez G, Brick C, Bricka T, Briker R, Brown AN, Brown GDA, van Aert RCM, Caldwell K, Capitan S, Capitán T, Chandler J, Charles T, Chartier CR, Chawdhary R, Cheng KJ, Chopik WJ, Clark B, Colvin VE, Comer CC, Costantini G, Coupé T, Cummins J, Czernatowicz-Kukuczka A, de Leeuw J, Dobolyi D, Druckman JN, Duan J, Dujmović M, Dunleavy DJ, Durkee PK, Emery C, Esterling KM, Evans TR, Fedor A, Fernández-Castilla B, Fiala N, Field JG, Fong N, Fonseca MA, Freeman ALJ, Freese J, Geiger SJ, Geng J, Getz LM, Geven LM, Gleibs IH, Gonzales DP, Gooty J, Gourdon-Kanhukamwe A, Greculescu C, Griffin SM, Grigoryan L, Grunow M, Gunby N, Hall B, Hanel PHP, Hannon EE, Harper S, Held MJ, Hickman L, Higgins NC, Hippel S, Hoeppner S, Hong S, Hostler TJ, Inzlicht M, Izydorczak K, Jaeger B, Jankowsky K, Jarke-Neuert J, Jensen M, Jokić B, Jolles D, Jolly P, Jones AM, Juanchich M, Kačmár P, Kapoor H, Keljanovic A, Koirala S, Kołczyńska M, Kouroupaki D, Kühnen U, Landgrave M, Larson MJ, Laulié L, Lawrence ACE, Le Forestier JM, Leahy KE, Lee S, Leslie J, Lewis SC, Limnios C, Lin H, Liu A-C, Lloyd JW, Ludvig EA, Lynott D, MacDonald J, Mallik P, Mallinson DJ, Marinazzo D, Martarelli CS, Matacotta J, McBride A, McHugh C, McMillan G, Méndez E, Metzger M, Michaelides MP, Michalak J, Micheli L, Miller JK, Milyavskaya M, Molden DC, Monjaras AG, Moreau D, Morrow A, Moya C, Mudrik L, Mulder LB, Munt KA, Nandi A, Nason K, Nast C, Nave G, Nax HH, Neubauer F, Nguyen PLL, Nichols AL, Nilsonne G, O’Boyle E, Oettinghaus J, Oh J, Oshana A, Ostermann T, Ostrowski RP, Oyebanjo A, Panczak R, Patrianakos J, Pavez I, Pavlov YG, Persson S, Perugini M, Peters K, Pieters C, Ponizovskiy V, Porter ND, Prenoveau JM, Purić D, Purol MF, Puthillam A, Quinn KA, Ramljak M, Reed WR, Ritchie M, Ritzau M, Roche SP, Rodela R, Röer JP, Ropovik I, Rothschild J, Saal J, Safadi H, Samaha J, Sanchez M, Sankaran S, Santos D, Sargent AC, Sauter M, Schmidt K, Schnabel L, Schroeder AN, Schuetz SW, Schuetze BA, Schulte-Mecklenbeck M, Schütz A, Sevigny EL, Shackleton E, Shafranek RM, Shaki S, Shakya S, Sirota M, Sisco MR, Sitnikov MM, Slevc LR, Smalarz L, Smith CT, Snyder JS, Sommet N, Sonmez F, Spellman BA, Stanulewicz-Buckley N, Stock G, Street CNH, Strømland E, Sundelin T, Syed M, Szabelska A, Szaszi B, Szumowska E, Tagat A, Täuber S, Tay L, Thapa S, Thatcher J, Tsaklakidou D, Tummers L, Turkovich E, Tutor MV, Urbanska K, van ’t Veer AE, van Assen M, van de Ven N, van den Goorbergh R, Vargo EJ, Vaughn LA, Vazire S, Vermeulen JM, Vo DTH, Volkman V, Wagenmakers E-J, Wagner D, Walasek L, Walter F, Warmelink L, Wei L, Weißflog MI, Weller N, Wichman AL, Wilbiks J, Williams JR, Wolfe K, Wort F, Wright R, Wulff JN, Xue X, Yan VX, Yang Y, Yoon S, Žeželj I, Zhang Y, Ziano I, Zogmaister C, Zupan Z, Zwaan RA, Nosek BA, Errington TM. 2026. Investigating the replicability of the social and behavioural sciences. Nature 652:143–150. DOI: https://doi.org/10.1038/s41586-025-10078-y

      Westlake H, David F, Tian Y, Krakovic K, Dolgikh A, Juravlev L, Bournonville TE de, Carboni A, Melcarne C, Shan T, Wang Y, Mu Y, Kotwal A, Pirko N, Boquete JP, Schüpfer F, Rommelaere S, Poidevin M, Liu Z, Kondo S, Ratnaparkhi GS, Chakrabarti S, Liu G, Masson F, Xiaoxue L, Hanson MA, Jiang H, Cara FD, Kurant E, Lemaitre B. 2026. Reproducibility of scientific claims in Drosophila immunity: A retrospective analysis of 400 publications. eLife 15. DOI: https://doi.org/10.7554/eLife.108404.1

      Youyou W, Yang Y, Uzzi B. 2023. A discipline-wide investigation of the replicability of Psychology papers over the past two decades. Proceedings of the National Academy of Sciences 120:e2208863120. DOI: https://doi.org/10.1073/pnas.2208863120

    1. eLife Assessment

      This valuable paper uses a mathematical model applied to a dataset of E coli / ESBL carriage and transmission to infer drivers of drug resistance in France. The strength of support for the study findings is incomplete. While the research question is of importance, and the mathematical model has structural and methodological integrity, numerous issues are noted: insufficient description of the data, lack of included equations and code, definitions of antibiotic use that are not complete, low sensitivity of assays for carriage, technical issues with statistical prior selection and parameter identification, and application of non-regional ECDC surveillance data to France.

    2. Reviewer #1 (Public review):

      Summary:

      The authors used a large dataset evaluating gut carriage of Enterobacterales and ESBL organisms from children aged 6-24 months as the basis for a modeling study to investigate what factors are most important for determining the prevalence of ESBL resistance. The modeling incorporated travel, a simple model of carriage duration (short and long), fitness cost of resistance on transmission and clearance, and antibiotic use. They found that antibiotic use is the primary driver of resistance prevalence, with transmissibility of resistant strains also important for setting the prevalence. Travel, while important when prevalence is very low, plays less of a role in maintaining prevalence once it is established (in keeping with other recent work). They estimated the fitness cost of resistance (terming a reduction of 14% on the rate of transmission and an increase of 23% on the rate of clearance as "low"). While the extent of assumptions and simplifications makes me skeptical of the quantitative conclusions, the qualitative ones seem reasonable and reinforce the long-held principles of the field--reducing antibiotic pressure and interrupting transmission--and highlight the importance of understanding the biological factors that shape the duration of carriage and the likelihood of colonization.

      Strengths:

      This study incorporates many of the factors that might influence the carriage prevalence of ESBL Enterobacterales. This builds on the work led by this group, both in primary data collection and in theory. Overall, it's such a tough problem that I commend the authors for trying to tackle it. The authors take a thoughtful, rigorous approach, acknowledging simplifications and assumptions where they need to, so as to evaluate the various factors shaping ESBL prevalence.

      Weaknesses:

      Part of the reason it's such a tough problem is that we have limited data to structure and parameterize a complex model.

      (1) The data are not sufficiently described.

      The primary data source for this modeling exercise comes from a study of 6-24-month-old children who underwent rectal swabs and evaluation of the carriage prevalence of Enterobacterales, and then whether these Enterobacterales were ESBL; moreover, the study included data on travel and on antibiotic use. Could the authors please direct us to these primary data? Could the authors also justify the parameters in their models from these data--for example, could they please provide the distribution of antibiotic use and the associated timing? Could they also explain why they decided to treat all Enterobacterales as if they were E. coli (line 307)? Is there evidence that all Enterobacterales occupy the same niche and compete with each other?

      (2) The model should be more fully described and the limitations explored/explained.

      - The authors should point to the code and the ODEs.<br /> - I understand the focus on the pediatric population; the authors argue that this is reasonable because ESBL colonization is similar across age groups. But presumably, antibiotic use differs across age groups, and there is colonization pressure from within households.<br /> - The authors only consider resistance to extended-spectrum beta-lactams and use of beta-lactam antibiotics, but ESBL Enterobacterales are often resistant to other antibiotics as well. How much does the use of other antibiotics also select for Enterbacterales that happen to carry ESBL resistance? "One bug/one drug" modeling, as done here, neglects the complexities of the actual patterns of resistance and range of antibiotic use.<br /> - Do the data support the T3 or S3 compartments, which, if I understand correctly, means no exposure to antibiotics can happen during three months after either treatment or travel? What do the data say about the patterns of antibiotic use? I'd imagine that the likelihood of antibiotic use is not homogenous, but instead, there are some who use repeated rounds of antibiotics.<br /> - Why do the authors exclude individuals who used antibiotics in the prior 7 days? What justifies that cutoff? The authors speculate that the impact of excluding these individuals is likely to be minimal; why exclude them, then? Did the authors evaluate the results if they were included?<br /> - What is the basis of "niche differentiation", as described starting on line 221? Why should clearance of one strain be slower when the strain co-occurs in a host with a strain of another type?

    3. Reviewer #2 (Public review):

      Overview:

      This study integrates several datasets into a unified modeling framework that incorporates several mechanisms thought to impact the spread of ESBL-resistant bacterial strains. The model accounts for tradeoffs between persistor and colonizer strains, travel rates, antibiotic treatment and strain clearance, direct competitive interactions, and, most importantly, a series of distinct costs associated with the carriage of ESBL resistance. The resulting 75-compartment model is internally consistent and structurally neutral. However, the parameter estimation is flawed in many ways, compromising the interpretations of the model.

      On the usage of the Swedish infant data set to estimate colonization and persistence:

      First, while other papers have taken similar approaches, the Swedish infant data set is fundamentally inadequate to estimate colonization and persistence rates. This is because very few colonies were typed per sampling event (2 to 6 colonies per event). The original authors themselves argued that strains of indistinguishable morphology would not be able to be differentiated by this method. They also provided data showing that strain identity was not directly related to colony morphology (same strain often displaying distinct morphologies).

      The consequence of this is that strains present in low abundance would be missed with a high likelihood. However, if they were to be stochastically sampled, this would count as a "colonization" event, and if they were missed in subsequent samplings, this would count as a "loss" event. In other words, the statistical methods described conflate within-host dynamics (which might lead to distinct within-host abundances) with between-host dynamics (colonization and loss).

      Beyond this conceptual issue, some technical aspects aren't particularly sound. The mean of the inferred posterior for the lambda and mu parameters are then used to calculate the beta, gamma, d, and epsilon parameters through a linear regression. The more technically correct way of doing this would be to directly infer these parameters from the data and obtain a full posterior for these parameters.

      This highlights another issue: these parameters are passed down to the next statistical model as point estimates, with no associated uncertainty. This artificially inflates the (already low) confidence of the estimates for the cost parameters.

      Finally, when this procedure generated parameters that were inconsistent with their expectations (clearance is too high to explain prevalence in France), they adjusted the parameters by discarding and recalculating their beta parameters to artificially enforce neutrality between their strains and enforce the expected prevalence. This is problematic because beta and gamma were jointly estimated, and there is no particular reason why some of them should be discarded. The more natural interpretation would be that parameters inferred from Swedish infants do not translate well to French adults, which should preclude their usage in this context.

      On the estimation of costs of ESBL resistance:

      The core of the second statistical model is to use prevalence data, travel data, and treatment data in conjunction with the previously inferred colonization and loss parameters to infer the costs of carrying antibiotic resistance. Therefore, the accuracy of this section is contingent on an accurate estimation of the previous parameters. However, these colonization and loss parameters are inherited with no uncertainty (just point estimates are passed down), which, as previously mentioned, generates an artificially precise posterior distribution for the resistance parameters.

      However, the most severe issue with the statistics lies in the choice of priors for the cost parameters. All of them are uniform in a positive range that implies a positive cost. Importantly, the average over a positive range will always be positive; therefore, this method will ALWAYS estimate a positive mean for the costs. Note that the posterior distribution of some cost parameters seems to peak around zero and abruptly decays with no mass to the left of zero. This is caused by the choice of prior. Had delta been allowed to be negative (i.e., antibiotic resistance carried a benefit, having the prior be uniform between -1 and 1), the posterior distribution would likely be much more symmetrical, and the confidence interval would have included 0.

      Restating, because the prior is a continuous function between 0 and 1, it contains infinitely more mass in the region that represents there being a cost (delta>0) than in the region representing no cost (delta=0). This means that it is a mathematical impossibility for this model to infer the absence of a cost.

      Therefore, the main finding of the paper ("We found that resistance is costly") is a mathematical artifact of the prior choice and of the model structure.

    4. Reviewer #3 (Public review):

      Cotto and colleagues integrated data analysis with mathematical modeling to examine extended-spectrum beta-lactamase (ESBL)-producing E. coli in France. While ESBL prevalence has risen globally, it has stabilized at approximately 6-8% across Europe. Established risk factors for ESBL carriage include prior antibiotic exposure and travel to high-prevalence regions, most notably South-East Asia. The dataset incorporated information on ESBL-producing E. coli and travel history in young children, and the model was calibrated to ECDC surveillance data on ESBL across Europe, supplemented by literature-derived parameters on antibiotic use, E. coli biology, and transmission dynamics. The authors report that ESBL-carrying strains exhibit a 14% fitness cost in community transmission relative to susceptible bacteria, yet are cleared 23% less frequently. ESBL carriage was strongly associated with factors that prolong gut colonization. Both antibiotic treatment rates and transmission efficiency were identified as key determinants of community-level ESBL prevalence.

      Strengths:

      The study addresses a clinically and epidemiologically important topic. The integrated modeling approach is methodologically sound and well-suited to disentangling the relative contributions of transmission and antibiotic selection pressure.

      Weaknesses:

      Several concerns regarding the data used in this study warrant consideration. First, model calibration relied on ECDC surveillance data pooled across multiple European countries, several of which have substantially lower antibiotic consumption than France (ECDC ESAC-Net Annual Epidemiological Report, 2024). Given that antibiotic use is a primary driver of ESBL selection, ESBL prevalence is likely to be heterogeneous across these settings. Calibrating to a geographically diverse dataset risks introducing systematic bias into parameter estimates that may not be representative of the French context. The authors should repeat the analysis using France-specific data, or, where this is not feasible, restrict the calibration dataset to countries with comparable antibiotic consumption profiles. Second, the travel exposure data may be insufficient to adequately capture importation dynamics from South-East Asia, as the cohort consisted exclusively of young children, a demographic less likely to travel to high-prevalence regions than older age groups. This may result in an underestimation of travel-associated importation as a contributor to community ESBL prevalence, and the generalizability of these findings to the broader population should be interpreted with caution.

    5. Author Response:

      We thank you for this assessment of our work and the positive assessment of the overall theoretical framework. We can fully answer the concerns, in particular regarding data quality, and will provide detailed answers in the following directions:

      “insufficient description of the data”: We will describe the data as much as possible and will share the data and the analysis code.

      “lack of included equations and code”:  We will share the mathematical equations in the supplementary material, and the full code on an online repository. The reason why our data repository (10.5281/zenodo.18480481) is not yet public is that it cannot be changed after publication. For the review process, we provide a github link to data and code here https://github.com/oliviercotto/eLife_epidR. We will ultimately share the link to the final version of the files on Zenodo.

      “definitions of antibiotic use that are not complete”: We will complete the definition of antibiotic use, which is the use of any antibiotic between 7 days and 3 months before sampling. Children who used any antibiotic 7 days before sampling were not included in the study. The type of antibiotic used is given in supplementary material S1: 93% of the antibiotics prescribed are beta-lactams (amoxicillin, amoxicillin/clavunalate, oral 3rd generation cephalosporins).

      “low sensitivity of assays for carriage”: the carriage study conducted in Sweden is used to get plausible estimates of carriage duration parameters in infants.

      • Strain definition is based mainly on randomly amplified polymorphic DNA (RAPD), not colony morphology. Strains with distinct morphology but the same RAPD profile are considered one strain. Conversely, it was checked that strains of the same timepoint with the same morphology most often had the same RAPD profile.

      • We did check that these data are not much affected by imperfect sampling: observations of a strain ‘disappearing’ from sampling then ‘reappearing’  at later timepoints are rare (14 out of 273 strains). This is why we did not correct these occurrences in the previous version of the analysis. In the revised version, we will add a description of these occurrences and correct them. This correction did not significantly alter the inferred parameters in our preliminary analyses.

      • At a broad level, the fact that E. coli clades vary in their carriage duration is very well established across multiple independent datasets; the precise value of carriage duration difference for “persistent” vs. “transient” that we inferred here (a two-fold difference, supplementary material S3) is actually relatively conservative, in the sense that other studies have detected more important differences. We will create a table summarising available evidence on colonization parameters of E. coli to show that the insights from the Swedish data are qualitatively robust.

      • Yet, we will conduct a range of sensitivity analyses to see how the inferred costs of ESBL resistance vary when varying differences in carriage durations, competition and niche differentiation.

      “technical issues with statistical prior selection and parameter identification”: We disagree there is a “technical issue” with prior selection: The fact that resistance is costly, hence that our priors are left-bounded at 0 for the cost parameters, is a prior expectation based on the observation that resistances do not go to fixation. If resistances only conferred an advantage in treatment, but zero cost, then they would quickly evolve to 100% frequency–contrary to what is observed in virtually all epidemiological studies of resistance. That said, we will relax the definition of these priors to test that the data is also compatible with a strong cost on some traits, and no cost or a “negative cost” on other traits.

      Regarding parameter identification and the specific comments on the inference of colonisation parameters: we re-inferred all colonisation parameters with direct inference assuming specific functional forms. This does not alter much the final colonisation parameters that we then use for our main inference. We will also conduct sensitivity analyses to examine how changing some of the colonisation parameters (carriage duration, competition and niche differentiation)  would alter the main inference.

      “application of non-regional ECDC surveillance data to France”: We will clarify our text, as there is a misunderstanding here: we do not use non-regional ECDC surveillance data for inference. We use ECDC data (i) for illustrative purposes, to show that trends in ESBL in France in this surveillance system are very similar to those observed in France in our focal dataset, thus showing the consistency and representativeness of our data. (ii) to give an overview of the weak and inconsistent association of ESBL with age across Europe, thus supporting the relevance of our approach even if our data concerns infants and children. We will make sure this is clarified in the updated version of the manuscript.

    1. eLife Assessment

      This study presents an important finding regarding the effect of Yoda molecules on PIEZO2 function, challenging the assumption that they selectively activate PIEZO1. The evidence supporting this claim is solid, but several methodological and conceptual issues need to be addressed. Overall, this work will be of broad interest to researchers working with PIEZO channels across various biological scales.

    2. Reviewer #1 (Public review):

      Summary:

      In this work, T. Wijerathne et al. investigated and reported the agonistic effect of Yoda1 and Yoda2 over PIEZO2 function using patch clamp electrophysiology, Ca2+ imaging, and molecular dynamics. They find that Yoda1 sensitizes PIEZO2 to membrane tension, can induce Ca2+ influx, and decreases its inactivation to a lesser degree than it does to PIEZO1 channels. Additionally, their data shows that Yoda2 sensitizes PIEZO2 channels to membrane indentation to a greater extent, but it has a weaker effect on channel inactivation than Yoda1. Interestingly, they report that a mutation in a conserved arginine between PIEZO channels can be used to abolish PIEZO1-mediated Ca2+ flux in response to Yoda molecules. As a whole, the results presented here should be put into perspective against previous and future works involving systems where both PIEZO1 and PIEZO2 might be expressed. This is especially true for works where Yoda1 has been used as a basis for determining the absence of PIEZO2.

      Strengths:

      The authors use multiple techniques to investigate how Yoda molecules affect the three most important biophysical aspects of PIEZO channels that, when changed, result in pathophysiological responses: a) sensitivity to mechanical stimuli, b) Ca2+ entry, and c) channel inactivation. Lastly, they find a specific amino acid/region that could be exploited for drug design and/or development.

      Weaknesses:

      The methods and discussion sections are lacking enough detail to fully evaluate the findings and put them into perspective, respectively.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript challenges the long-standing assumption that Yoda1 and Yoda2 are PIEZO1-selective activators. Using patch-clamp electrophysiology and calcium imaging in HEK293TΔPZ1 cells overexpressing PIEZO2, the authors demonstrate that Yoda1 potentiates PIEZO2 stretch-activated currents to a similar extent as PIEZO1 and slows PIEZO2 poking-current inactivation (albeit with lower efficacy). They further show that the more potent analog Yoda2 affects PIEZO2 at nanomolar concentrations and use mutagenesis and molecular dynamics simulations to propose that Yoda2's benzoic acid group forms a transient salt bridge with R1724 in the putative Yoda binding pocket, explaining its enhanced potency.

      Strengths:

      The authors are established Piezo/biophysics experts; the study is highly important, technically competent, and carries significant implications for the reinterpretation of prior work that used Yoda compounds as PIEZO1-selective probes.

      The core finding that Yoda1 modulates PIEZO2 stretch currents is convincing and important. However, several conceptual, methodological, and presentational issues need to be addressed before acceptance, as detailed below.

      Weaknesses:

      (1) The abstract states that Yoda1 potentiates PIEZO2 "as efficaciously as PIEZO1." This claim is accurate only for stretch currents and single-channel open probability, but the paper itself demonstrates important asymmetries: i) Yoda molecules slow PIEZO2 poking-current inactivation ~2-fold, versus ~5-10 fold for PIEZO1 (Figure 3b and ref #60). ii) Spontaneous Ca²⁺ entry via PIEZO2 requires non-physiological conditions (high extracellular Ca²⁺, hypertonic solutions) that are unlikely to occur in native cells.

      The abstract should be revised to clearly qualify where equivalence holds and where efficacy differences exist. IMO, the current wording risks overcorrecting the historical bias (PIEZO1-only) by going too far in the other direction.

      (2) Related concern: the PIEZO2 Ca²⁺ signal in Figure 2 is only detectable using a Ca²⁺-boosted solution (CBS ie 30 mM Ca²⁺). Physiological extracellular Ca²⁺ and cells normally do not experience sustained hypertonicity at these magnitudes. The authors should explicitly clarify that the practical implication of their findings is primarily for electrophysiological (patch-clamp) experiments and that the Ca²⁺ imaging caveat applies only under amplified conditions. Specifically, the authors should state that in standard Ca²⁺ imaging assays with physiological buffers, PIEZO2 is unlikely to confound Yoda1 results.

      Related point: Can cytochalasin D (CytoD) restore a Yoda1-dependent Ca²⁺ signal in physiological saline? This would help determine whether the weak PIEZO2 response is primarily a membrane tension issue (cytoskeletal tethering) versus intrinsically lower channel expression or permeability. The authors already have tagged PIEZO1/2 constructs and could, in principle, normalize by surface expression.

      (3) The mean inactivation tau values for wild-type PIEZO2 poking currents in both DMSO and Yoda1 conditions (Figure 3b, approximately 15-40 ms range) appear substantially higher than values reported in published literature (typically 5-10 ms; eg, PMID: 20813920). This discrepancy needs to be addressed.

      (4) The authors perform all MD simulations on a truncated PIEZO1 model and justify this choice by noting that the Yoda binding region is highly conserved between homologs. This is a reasonable and defensible starting point given the availability of well-validated PIEZO1 simulation set ups in their lab. A few points are nonetheless worth addressing: While PIEZO2 simulations are not strictly required, the authors are encouraged to briefly discuss whether any long-range structural differences between PIEZO1 and PIEZO2 (outside the binding site itself) could influence Yoda2 binding dynamics, particularly in light of the chimera data showing that PIEZO2 sequence in repeat A abolishes Yoda1 sensitivity. This reviewer still doesn't understand the reason behind this discrepancy despite it being acknowledged in the text.

      Another MD-related comment is that three simulation replicas (which is impressive for such a big system) show markedly different salt bridge occupancy (82.6%, 49.7%, 99.8%; stated in the text). This wide variation suggests incomplete sampling in at least one replica. The authors should provide RMSD plots for ligand and protein backbone to assess convergence and possibly discuss whether the 49.7% replica represents a genuinely distinct binding mode or incomplete equilibration.

      (5) The Discussion proposes that PIEZO2's weaker Ca²⁺ response to Yoda1 could partly reflect lower membrane expression. Since the authors already have fluorescently tagged PIEZO1 and 2 constructs, a simple fluorescence intensity comparison between the two (acknowledging it would reflect total rather than surface expression) could provide at least indirect support for this claim. Alternatively, if such a comparison is not feasible, the authors may consider removing membrane expression from the list of proposed explanations or explicitly acknowledging that this remains unsubstantiated speculation. The max poking currents may somewhat and roughly indicate the level expression difference too, if done exactly side by side.

      (6) The abstract or concluding remarks should highlight that Dooku1 is not PIEZO1-selective in its agonist-like action on PIEZO2, and that Cmpd15/Cmpd64 appear to be better PIEZO1-selective tools. This nuance is buried in the Results section.

      (7) The authors should not cite PMID 31015490. Clearly, any work on MCC13 is confounded by the overwhelming expression of PIEZO1 (PMID: 42084270). Instead, the authors should also cite the literature from others who have clearly recorded stretch currents from PIEZO2 before the cited studies (eg, PMID: 37590348).

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript reports that Yoda1 and Yoda2 agonize PIEZO2 in a manner similar to PIEZO1, increasing open probability and stretch sensitivity, but the mechanism underlying this sensitivity is incomplete. Mutagenesis was shown exclusively in PIEZO1, with no corresponding mutagenesis in PIEZO2, so the proposed mechanism in PIEZO2 is inferred by homology rather than directly tested. All experiments use mouse PIEZO2, and the human ortholog should be used before generalizing the proposed reinterpretation of the field.

      Strengths:

      The pressure-clamp electrophysiology demonstrating a shift in half-activation pressure for PIEZO2 is compelling evidence in support of the central claim.

      Weaknesses:

      (1) In the single-channel recordings (Figure 1a), it's unclear how many channels were present in those patches. After applying -60 mmHg pressure, multiple channels would be activated (as seen in Figure 1e). The number of channels in the patch and their inactivation rate could significantly influence the open probability in such experiments. To overcome this, in the original Yoda1 article (Syeda, Ruhma, et al. eLife 2015), no additional pressure was used. Additionally, the reported open probability comparison (n=7 Yoda1 vs n=17 DMSO patches) has an SEM nearly as large as the effect itself (0.30 {plus minus} 0.11), consistent with a small number of outliers driving this. The underlying mean open and shut times are reported without any statistical test; only the derived open probability receives a p-value. Additionally, in Figure 1a, the Yoda1 condition noise is different from the control. This should be stated if noise filtering was applied and how, given that this could affect open probability analysis.

      (2) The calcium imaging data in Figure 2 raise significant concerns regarding the chemical activation claim. The calcium-boosted solution (30 mM Ca2+) is not physiological and appears to be generally stressing cells rather than specifically activating PIEZO2: the control condition under CBS already shows an elevated signal, consistent with cells being unwell at this calcium concentration, and adding Yoda1 on top of this shifted baseline raises further questions about specificity rather than confirming it. Separately, it is unclear why DMSO alone produces measurable PIEZO2-associated calcium influx in HBSS, a result that is not addressed in the text. Figure 2 should clearly indicate when DMSO/Yoda1 perfusion was initiated, and y-axis labels are missing from panels A and B.

      (3) In the poke experiments, an activation threshold should be calculated and reported, and amplitude data (e.g., peak current versus indentation depth) should be shown rather than only inactivation tau values. It is also unclear why mClover3- and N-GFP-tagged constructs were used in these experiments, since electrophysiological recording already confirms channel expression without requiring a fluorescent tag.

      (4) For inactivation kinetics (Figure 3b), the authors use unpaired comparisons across separate cells, whereas the deactivation experiments (Figure 3c) use paired; it should be applied to the inactivation experiments as well. Deactivation kinetics for PIEZO2 itself should be shown. If the claim is that Yoda1 acts on PIEZO2 through the same mechanism proposed for PIEZO1, then a PIEZO1/2 chimera should be expected to show a corresponding effect on deactivation tau; instead, this chimera is reported as completely Yoda1-insensitive despite both parental channels being Yoda1-sensitive, as shown in this study.

      (5) Given that this reflects a different experimental paradigm for Yoda EC50, PIEZO1 should be included within Figure 4b. Additionally, EC50 bar plots should be present on this figure. The inactivation time constant for PIEZO2 without Yoda1 is inconsistent across figures, below 20 ms in Figure 3b but above 20 ms in Figure 4c.

      (6) Finally, the modeling is performed exclusively on PIEZO1, whereas the manuscript's central focus is PIEZO2. It is therefore unclear whether the proposed structural mechanism, including the basis for Yoda2's reduced efficacy on PIEZO2, can be directly extrapolated to PIEZO2.

    1. eLife Assessment

      In this manuscript, the authors describe a new member of the KCNE auxiliary subunits of potassium channels from a lamprey. This new subunit represents an early evolutionary member which confers new properties when expressed along with KCNQ channels. The authors present convincing evidence from several experimental approaches. The contents of this manuscript are important and should be relevant to understanding both the mechanism of modulation of KCNQ channels by KCNE subunits and the evolutionary history of these subunits, which this manuscript now extends to the divergence of early vertebrates.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

      Weaknesses:

      (1) No error bars visible for lamprey Q1 isoforms (open symbols) in Figure 2G. No statistical comparison was provided to indicate whether lamprey Q1 isoform V1/2s are significantly different (nor in Supplementary Table 1).

      (2) There is the same issue in Figures 3 and 4. No appropriate statistical comparison is made between V1/2s for different truncations of PmKCNE0 (Figure 3), or between KCNQ1 species isoforms with and without PmE0.

    3. Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

    1. eLife Assessment

      In this important study, Boudjema et al. use cell culture models and high quality advanced microscopic imaging to provide detailed analyses of the cellular processes underlying centriole amplification, apical migration, and assembly of hundreds of motile cilia in multi-ciliated cells. The authors present convincing evidence showing that in these cells all the molecular and cellular steps controlling centriole biogenesis that in cycling cells extend over almost two cell cycles, occur within a single cell cycle variant. This work provides a better understanding of the regulation and order of these processes and is of interest to all cell biologists and in particular researchers studying centrioles and cilia.

    2. Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      Comments on revised version.

      The authors have significantly improved the manuscript, by refocusing it, introducing text and figure changes, and by adding new data including functional analyses. The revised version now has convincing data that support the claims. All my remaining concerns have been addressed.

    3. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the Cen2-GFP; mRuby-Deup1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described Cen2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of Deup1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells(treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules and microtubule-based transport in different stages of differentiation in brain MCCs. Addressing the role of microtubules during different stages of centriole amplification required development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative spatiotemporal analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive and are further strengthened in the revised version through additional analysis and adoption of new methods. This comprehensive analysis advances our understanding of MCC biology regarding the involvement of microtubules.

      Comments on revised version.

      The revised manuscript is substantially improved, and given the scope, it is appropriate that it primarily establishes a detailed spatiotemporal framework. That said, a few points would further strengthen clarity and impact. First, several observations naturally raise follow-up mechanistic questions, for example whether additional cytoskeletal systems such as actin contribute to steps like centriole apical migration. A slightly more detailed framing of these open questions would help guide future work. Second, some terminology introduced to label observed microtubule-based structures (for example "nest") may not be essential. Finally, while the authors have increased quantification, some analyses would benefit from super plot-style displays with replicate-level comparisons, particularly for intensity-based readouts.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully addressed the insightful comments provided by the reviewers which thoroughly increased our comprehension of the dynamics of centriole amplification. The manuscript has been revised accordingly and put in the context of the two papers we published since our last submission, showing that MCC differentiation is a genuine cell cycle variant. A point by point answer to all reviewer comments is provided below.

      Briefly:

      We have streamlined terminology and nomenclature in text and figures / better define experimental conditions with nocodazole

      We have tested the role of dyneins in the dynamics of centriole amplification

      We have done correlative light and electron microscopy on the early stages of centriole amplification

      We have analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors

      Collectively, this allowed us to make a clearer parallel with what occurs during centriole duplication and to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles.

      We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Public Reviews:

      Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      We have streamlined terminology and nomenclature, clarified the description of the complex events, and test the role of dyneins in centriole amplification. Microtubules density in MCC does not allow to extract information from imaging. In addition, we have done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Altogether, our new data allowed to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles. We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Reviewer #2 (Public Review):

      This important work will be of interest to centriole and cilia cell biologists. It describes in detail how microtubules control multiple aspects of centriole amplification in brain multiciliated cells. This study provides a greater time-resolved and molecular proteomic mapping of the different steps involved, with or without microtubule disruption. Boudjema et al. show that microtubules are important throughout the centriole amplification process, from the early stages, where the procentrioles emerge from a pericentriolar "nest", through the growth stage where microtubules maintain the perinuclear localisation, to the detachment stage, where microtubules assist in perinuclear disengagement and apical migration. The results are generally well supported by the evidence, but the manuscript would benefit significantly from some heavy editing to introduce more niche terms, standardize abbreviations in text, and labels on figures to help bring the readers, especially non-specialists, along with them - increasing the accessibility of their work.

      We thank the reviewer for his/her enthusiasm. We have streamlined terminology and nomenclature and clarified the description of the complex events to increase the accessibility of our work. We also replied points by points to his/her specific comments.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical-basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, and MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the CEN2-GFP; mRuby-DEUP1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described CEN2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of DEUP1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM, and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells (treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules in centriole amplification. Addressing the role of microtubules during different stages of centriole amplification required the development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive. This comprehensive analysis advances our understanding of MCC biology.

      Weaknesses:

      The role of microtubules and other molecular players during different stages of centriole amplification in brain MCCs can be further studied and strengthened using the tools developed in the manuscript. A more quantitative description of some of the analysis performed in the manuscript is required to strengthen the conclusions.

      We thank the reviewer for his/her enthusiasm. We have tested the role of dyneins in the dynamics of centriole amplification, done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Recommendations for the authors:

      As you will see, all reviewers felt that the analyses of the involvement of microtubules should be strengthened by including controls and additional experiments. Also, they agree that significant text editing would help to improve the manuscript's accessibility and readability.

      Specifically, they would suggest (1) streamline terminology and nomenclature in text and figures; (2) better define experimental conditions with nocodazole (concentrations used, effect on microtubules, effect on canonical centriole duplication); and (3), in the absence of other complementary genetic perturbation experiments, add a limitations paragraph in the discussion about conclusions drawn from nocodazole treatment alone.

      Reviewer #1 (Recommendations For The Authors):

      Main issues:

      (1) The authors use variable terminology to describe the same or similar events/structures. For example, in Figure 1 they refer to "centrosome stage" where they observe a pericentrin "cloud", which they later refer to as a "nest". In all other figures the first stage is not referred to as the "centrosome stage" but as the "cloud stage". Again, they also describe the "cloud" as a "nest" occasionally, but not always. In the cartoon, the nest is termed "centrosome cradle". The variable and inconsistent use of terms is confusing and the authors do not provide any explanation for the use of one vs. another.

      The text is now corrected. The centrosome stage corresponds to the stage preceding the beginning of centriole amplification in MCC progenitor. The pericentrosomal cloud of centriole and deuterosome elements forms later on, during the amplification A-stage. The formation of this cloud marks the beginning of A-stage, and persists up to G-stage where it dissolves. When we show that the cloud hosts the first stages of centriole biogenesis, we defined it as a “nest”. We do not use anymore the term craddle.

      (2) What prompted the authors to use the term "nest"? It gives the impression that they describe aspecific physical entity/structure (also depicted in this way in Figure 3P, with microtubules outside of this structure), but what is the evidence for this?

      The cloud is the spatial entity and the term “nest” is used to define a function of this transient compartment. We decided to keep the term “nest” as we now identified it with correlative light and electron microscopy, in addition to U-ExM, and show that the accumulation of centriole and deuterosome elements is accompanied by the formation of immature procentrioles, deprived of MT walls, as well as immature and empty deuterosomes. The scheme with MT outside the cloud/nest is misleading as we see MT organized by the mother centriole. We have now changed this.

      (3) The "nest" may simply be a dynamic accumulation of precursor particles around the centrosome, similar to what has been described for centriolar satellites. Rather than proposing a new entity, I suggest testing whether the "nest" particles may colocalize with PCM1 and thus may be related to centriolar satellites. Based on the data, the nest would simply be the centrosomal MTOC that organizes a radial microtubule array on which particles move around its center. In the absence of other evidence, I am not convinced that a new term is needed.

      We totally agree with the reviewer: the centrosome, as MTOC, concentrates centriolar and deuterosome components. This cloud is consistently dissolved when MT are depolymerized or dyneins inhibited. So, the physical entity is a “cloud”. We used the term “nest” to propose one function for this cloud which is to form deuterosomes and centrioles, before they move away for maturation. In fact, deuterosome and centriole formation are hindered when the cloud is dissolved. We have tried to edit the text all over the manuscript to make it clearer.

      (4) Role of MTs: are microtubules required or do they just facilitate some of the investigated events?

      The reason why the role of MT has not been tested yet during centriole amplification is probably because MT not only constitute the cell cytoskeleton on which molecular motors ride to transport cargos or distribute forces, they are also the core component of the structures we are studying. This is why we have tested a range of nocodazole concentrations and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (Fig. 4 Supplementary 1A-B). This may lead to an underestimation of the role of MT but we cannot study the role of MT on centriole amplification if centrioles cannot be formed.

      Does multi-ciliation in these models eventually occur normally under the concentrations and treatment conditions used here? This should be tested and discussed in the context of whether microtubules are indeed required and at what step of the entire process (amplification, migration, ciliogenesis) they may be critical.

      We did both chronic and acute treatments.

      Chronic treatments were done to test the overall efficiency of centriole amplification when MT (or dyneins) are perturbed. Chronic treatments were used to assess the role of MT (or dyneins) on the global efficiency of centriole and deuterosome formation (number of cells able to amplify, number/size/loading of deuterosomes, final number of centrioles (Fig. 4H-I, Fig. 4 Supplementary 2 B-D). In these chronic treatment, we focused on centriole amplification and not ciliation since it was the scope of this study. Also, we did not take ciliation as a readout of amplification because ciliation is relying on MT polymerization.

      Then, we also did acute treatments to test the role of MT (or dyneins) at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration; Fig. 4, 5, 7, 8 and associated supplementary figures). Since one stage is dependent on the precedent one, this enabled us to decipher the direct role of MT (or dyneins) on each single stage. We have now edited text, methods, legends and pictograms to be clear on whether acute or chronic treatment was done.

      (5) Can the authors include control (non-amplifying) progenitors in their analyses? It would be useful to know what the signal and distribution of each specific marker are before differentiation begins (before the cloud stage).

      Non amplifying progenitors are analyzed and constitute the so-called “centrosome stage”. We have now precised it and called it the “progenitor stage”.

      (6) Figure 2: Again, the terminology is confusing, since the authors describe that DEUP1 forms a "cloud" with centrin during the A stage.

      Corrections have been done as explained in point 1.

      (7) Description Figure 3: the authors introduce yet another term: "halo" A-stage. Is this the early A stage? Again, this is not explained and confusing. More systematic and consistent description is needed.

      Corrections have been done as explained in point 1. The term halos is used un the lab as it was the first term we used in our Nature paper in 2014 in reference to the halo described by Erich Nigg when they overexpressed Plk4. It was an error to use it in the manuscript.

      (8) Nocodazole treatments: the used concentrations are quite high.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely (Fig. 4 Supplementary 1A-B).

      (a) To avoid non-specific effects the authors should test what the minimal concentration is that completely depolymerizes microtubules in their cell model and perform analyses at this concentration.

      We have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A-B), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). In case it was not clear, we refer to this now several time and more clearly in the text and methods.

      (b) They should demonstrate depolymerization of microtubules by microtubule staining in the acute and chronic noc treatments and at the different noc concentrations used.

      This is, and was, in supplementary material (same, Fig. 4 supplementary 1A).

      (c) The authors should demonstrate that the used nocodazole concentrations do not impair normal centriole biogenesis during the cell cycle in these cells; if so, impaired assembly of centriole wall MTs may contribute to the observed effects in Figure 4.

      As mentioned in point 8b, we have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). The ability of the cells to form centrioles during chronic treatments were always assessed using immunostainings of SAS6 and/or CEN2-GFP signals (now exemplified in Fig. 4 Supplementary 1B). We also did EM analysis on cells treated with the highest doses of nocodazole (Nocodazole 10 uM for 24h) and this showed that centrioles can form with, what seems to be MT walls, in cells totally deprived of cytoplasmic MT fibers (Fig. 4 Supplementary 3-4). However, this does not show that all the cells can, because the number of cells that can be analyzed by EM are not sufficient to conclude. Also, one cannot assess whether MT walls are properly polymerized. However, the absence of MT walls should not change the results of the Figure 4, which are based on DEUP1, SAS6 or CEN2-GFP signals for deuterosomes and centrioles. Also MT depolymerization affects the formation of deuterosomes, which should not be altered by MT wall defects as it is not affected, even when centriole formation is blocked (LoMastro et al., 2024). Last but not least, we now show that blocking dyneins, as a comparable and even greater effect, on the formation of the cloud, deuterosomes and centrioles (Fig. 4C-I and Supplementary Fig. 4), which confirms that MTOC function, rather that MT wall formation, explain the centriole biogenesis alteration shown in Figure 4.

      (9) The authors repeatedly refer to the centriole-to-centrosome conversion of amplified centrioles and how this resembles centriole-to-centrosome conversion during the cell cycle. However, they incorrectly claim that this occurs at the G2/M transition. PLK1-dependent modification occurs at this stage, but conversion and PCM recruitment only occur after mitosis (see original work by the Tsou lab, which needs to be cited here).

      We agree with the reviewer. We have now added additional data to show clearly that centriole biogenesis, which requires two cell cycles to proceed in cycling cells, is accelerated during the MCC cell cycle variant where the elongation and maturation cycles are superimposed. This is now clearly shown in Fig. 3, 5, 9 and discussed.

      (10) Figure 6H-J: the authors claim that at low noc concentration, more D-stage cells showed incomplete disengagement than in controls, but the effect is shown only for the highest 10 µM concentration. Do any eof the phenotypes in Figure 6 also occur at the lowest noc concentration (assuming it depolymerizes MTs)? Again, it is crucial to demonstrate this, to exclude unspecific effects not linked to MT depolymerization.

      An error was made on the figure (but not in the legend). In Figure 6, chronic treatments are at 1 or 5 µM. Only acute treatments were done using 10 µM. In both cases, MT are not entirely depolymerized in these experiments (Fig. 4 supplementary 1A).

      (11) Disengagement, Figure 7: The authors describe that DEUP1 signal spreads all over the cytoplasm and becomes diffuse during this process, but one cannot see a diffusive signal throughout cells in the figures.

      We pushed the contrast to make it clearer but the deuterosomes are still bright at this stage and it is difficult to have both signal clear (now in Fig. 6B). We have also changed the example in video (now video 19) to show it more clearly with DEUP1 channel alone.

      (12) Figure 7: localization of disengaged centrioles at microtubule "nodes" is not clear from the images. There are many centrioles and random colocalization may be expected simply based on the high number. Higher resolution and/or magnification and quantification would be needed.

      We have edited and now say that centrioles “colocalize” with MT which, since centrioles nucleate MT, seems normal. We agree that it could be random, but given the density of MT, and the number of centrioles, it does not seem opportune to us to quantify. We can just say that we never see centrioles is regions that are deprived of MT.

      (13) The term "diffusive" to describe slow centriole movements in Figure 8 suggests that it is not motor or force-dependent, but there is no evidence for that. Movement based on opposing forces could produce a similar result, but would not be considered diffusive.

      We agree. We have changed “diffusive” by “diffusive-like”.

      (14) The manuscript would greatly benefit from the analysis of some candidate motor activities that may drive the movement and migrations of centrioles in this system. This would support the importance of the microtubule network for the specific steps in these processes, and better define its role beyond "being required". Dynein may be a candidate or minus end-directed kinesins. Since chemical inhibitors are available, these types of experiments would be straightforward.

      We formerly tested ciliobrevin but had hard time because of the small stability of the drug. Since our submission to eLife, we tested dynapyrazol and dynarestin and found dynapyrazol very efficient in dissolving the Golgi, a good readout of dynein inhibition. We sought to test the role of dyneins, using dynapyrazol, on (i) the formation of the pericentrosomal cloud in A-stage, (ii) the oscillation of DEUP1+ structures during A-stage, (iii) the number, size, loading of deuterosome, (iv) the final number of centrioles, (v) the migration to the nuclear membrane and (vi) the final apical migration of centrioles. The results are now inserted in main and associated Fig. 4, 5, 7, 8, 9.

      (15) Discussion:

      "the role microtubules" lacks "of"

      This is now edited.

      "This lack is..." Lack of what?

      This is now edited.

      "reflexive link" - meaning of "reflexive" is not clear in this context

      We have removed it.

      In my opinion, the study does not identify a nest composed of DEUP1, PCNT, and Centrin2; it only shows that these components accumulate as particles around the centrosome, which functions as MTOC. Consequently, it seems that the "nest" does not exist when MT is depolymerized. One could consider the center of the centrosomal MT array as a nest in this context, but there is no evidence of a specific new structure as suggested by the way the term is used in the manuscript.

      This is what we want to say: the center of the MT array become a nest in this context. We do not state that there is a specific new structure. We just say that MT and dynein dependent concentration of centriole and deuterosome components exists and that this region nests the birth of centrioles and deuterosomes. Also, this compartment is restricted in time and space, which justifies to use a specific term. The MTOC exists in the progenitor cell, while this compartment, marked by DEUP1, Centrin, PCNT accumulation, appears at the beginning of amplification and grows during A-stage to be dissolved at G-stage when all the deuterosomes and centrioles have moved away.

      What is the evidence that "DEUP1 is a centrosomal protein before building deuterosome structures"? It would be good to refer to the specific experiment. Does DEUP1 localize at centrioles also in the absence of microtubules? If not, I would not consider it a centrosomal protein.

      We have removed this statement to avoid misinterpretation.

      "This reminds the centriole-to-centrosome conversion..." the sentence is missing an "of"; also, again the authors confuse the order of events during the cell cycle, where centrosome conversion occurs after completion of mitosis, not at G2/M transition.

      We have removed this statement to avoid misinterpretation. Also, see Point 9.

      "microtubule dependent nuclear migration" should be rephrased; it sounds as if the nucleus migrates.

      This has been changed

      The following discussion of disengagement being linked to association with the nuclear envelope and resembling the process in cycling cells is misleading. In cycling cells movement of centrioles along the nuclear envelope occurs at G2/M and drives centrosome separation (separation of centriole pairs) in preparation for mitosis, not centriole disengagement.

      We are now clearer. We compare centriole-loaded deuterosome organization around the nuclear membrane to the migration of new centrosomes during early prophase (Fig. 5F-H, Fig. 5 Supplementary 2G-K).

      Regarding the possibility that forces by microtubules generated by the daughter centriole drive disengagement also in cycling cells, I would argue that this is unlikely since the daughter centriole can only nucleate microtubules after disengagement has occurred (and conversion to centrosome/PCM recruitment). Once this happens, it may physically separate the disengaged centrioles, which is a different type of activity. Indeed, originally the term "disengagement" was coined to specifically describe the loss of the perpendicular engagement of daughter centrioles with their mothers (Tsou and Stearns, Nature, 2006).

      We have removed this statement to avoid misinterpretation. The perpendicular engagement is difficult to assess on deuterosomes but we do see by live imaging, that attachment changes during D-stage, before centrioles detach clearly from deuterosomes.

      "high resolutive" should be "high resolution"

      Edit done.

      "splitted" should be "split"

      Edit done.

      "Consistently, when the mitotic oscillator is dis-inhibited and cells enter pseudo-mitotic events, centrioles show clear and rapid cell-cycle like clustering" This sentence is not understandable without further explanation; what does mitotic oscillator refer to? What are pseudo-mitotic events? What is cell cycle-like clustering?

      We have removed this statement.

      Minor:

      (1) Abstract: "Centriole number must be restricted to two..." Since cells are born with two centrioles and have 4 centrioles (2 pairs) when they enter mitosis, this sentence is inaccurate.

      The sentence has changed.

      (2) Abstract: "reflexive link"; I am not sure what the term "reflexive" refers to?

      We have removed this statement to avoid misinterpretation.

      (3) Figure 1C, D: it should be described better that the larger magnification panels represent overlays of many cells and what marker they show. This is not obvious since the smaller single-cell panels always show two different markers. Also, it would be more useful to show also single cells in the magnified view. The overlay does not allow us to see if a marker forms a cloud or a single dot, which is as important as the cell-to-cell variation in distribution.

      We have clarified this in the text and the legend. The cell-to-cell variation cannot be estimated with the overlay, but the projection from several cells (number precised) allows to see that the signal is confined in a restricted region. Or not. Which is what we wanted to analyze.

      Related to the above, the authors say that pericentrin forms a cloud at the top left in panel D, but there is only one confined centrosomal dot in the single-cell panel.

      The sentence has changed.

      (4) Results, Figure 2F; video 4: The authors claim connection and disconnection of DEUP1 aggregates with centrosomal centrioles; can the authors comment on the spatial resolution including in z in this movie to support this claim? Can they exclude that the structures are in proximity of each other rather than "connected"?

      This is a single z-section of 500nm. The resolution in xy is 128nm/pixel. Given the sizes of deuterosomes and a mature centriole, and given the fact that we observed this dynamics in several cells in live, we can state that the structures are connected. This is consistent with deuterosomes frequently observed “kissing” the daughter centriole by EM in the present manuscript (Fig. 2D, Fig. 2 supplementary 3 and 4 and Fig. 4 Supplementary 3-4). One has to look carefully at the daughter centriole (marked “dc”) and span in on the serial sections to see the connected deuterosome (marked by a star): this is at very early stage and therefore it is small. We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both structures are frequently sticked to each others on tens of nanometers.

      (5) The term "dynamics" as used in the manuscript should be plural.

      It has been used plural, except when for “dynamic microtubules” and “dynamic attachment to the nucleus”, which we think is ok? We have not found any other singular uses in our manuscript.

      (6) Figure 5: what does "YL1/2 procentriole intensity" refer to in panel F? This should be the intensity of microtubule asters.

      This has been modified.

      (7) Figure 6 - supplement 1B: contrary to the claim in the text, one cannot see tight colocalization with the nuclear pore marker. This seems to be a very small subset of particles and even in those cases colocalization is not tight. Also, what is the relevance of nuclear pore colocalization?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. What we want to say is that there is a tight connection with the nuclear envelope as shown by the localization of NPC on the same z-section as centrioles. This is why we present a single z, to show that centrioles and NPC are on the same z-plane of 500nm. NPC are stained to outline the nuclear membrane. This is also clearly visible for G-stage centrioles in the XY plane. We have now added an entire z-stack on video 18.

      Reviewer #2 (Recommendations For The Authors):

      To improve accessibility of their manuscript, we would suggest making the following edits:

      (1) Define 'specialist' or 'niche' terms each time you introduce them, such as 'pericentrosomal nest', or 'flower-like structures'.

      This has been clarified.

      (2) Have a think about abbreviations, again ones that work for people outside the project- this paper uses 'PC' for 'procentriole' but for many 'PC' is 'Parental centriole' or Figure 6J talks about 'D total' or 'D partial', leaves readers confused.

      This has been clarified.

      (3) Standardize your abbreviations throughout particularly for your treatments- sometimes Noco sometimes, NOCO, or your imaging experiments sometimes Cen-GFP, sometime CEN2-GFP (Figure 7A, D vs. Figure 6) or DEUP1- mRuby, DEUP1-mRuby3 or mRuby3-DEUP1?

      We now use Nocodazole or Noco in the text and the figure respectively, CEN2-GFP and mRubyDEUP1.

      (4) About 10% of the population, including several key figures in this field, are red-green color blind. Although 4 colour fluorescence is difficult to get right for everyone, choosing palettes (especially for two colour panels) is inclusive. More so, greyscale or inverted monochrome images make it easier for everyone to visualize changes in localization, size, and intensity. Red on black small foci is particularly difficult to discern. For example, Figure 3 - more individual channels in grayscale with arrows to mc, dc, and cilia would be helpful - difficult to distinguish stainings.

      We thank the reviewer for this comment and for this recommendation of being more inclusive. We have done the changes.

      To improve the conclusions drawn, we suggest some revisions below:

      (1) Since the paper really hangs on it, a clearer description of the rationale for when, how long and how much nocodazole treatment was done is needed. The logic currently is difficult to follow seemingly random jumps 10x concentration are used. Microtubules control many aspects of cell biology and could be impacted. For example, I particularly found Figures 6D and H difficult to follow i.e. the timing for 6H seems off.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely. We have of course tested a range of nocodazole concentrations at the beginning of the study and shown the extent of MT depolymerization under each treatment. We used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4 reviewer 1). The level of perturbation of MT and consequences on centriole formation at the different timings and doses were done for each experiment and are exemplified in Fig. 4 supplementary 1A-B. This figure was already present in the first version of the manuscript but we have now edited text, methods and pictograms to clarify this.

      (2) Perhaps an extension of this point- in general how interdependent are the processes? If there is a defect at the nest stage, how much are the later defects secondary to this, or do MTs genuinely play direct roles at all stages or are these knock-on effects? How do the authors rule this out? Defects in the nest, lead to smaller and more DEUP1+ foci, with defects in concentrating procentriole factors and centrin, which lead to... For example, Figure 4B looks like centrin is reduced upon noco treatment? Does noco treatment affect Cetn2GFP levels globally? Individual channels grayscale would help visualise this better.

      See also our answer to reviewer 1 point 8c.

      The stages are indeed interdependent. This is why we did both chronic and acute treatments. Chronic treatments were done to test the overall efficiency of centriole amplification when MT are perturbed. We typically used low dose of 1µM because nocodazole remains 48h in the culture medium. Acute treatments were done to test the role of MT at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration). Most of the acute treatments were done live and nocodazole was applied after the first time point of live monitoring. We used 10µM to have a rapid effect, and because nocodazole remains only several hours in the culture medium. This allowed to monitor the stage “n”, in cells where the stage “n-1” was completed without any drug which allowed to analyze a stage without having perturbed the precedent one.

      We now also test the consequences of dynein inhibition using both acute and chronic dynapyrazole treatments. We show that except for centriole migration, dynein inhibition phenocopies MT depolymerization (centriole number, perinuclear organization and disengagement as well as deuterosome number/loading/size).

      Nocodazole chronic treatments do affect intensity of CEN2-GFP at G-stage centrioles suggesting an altered A-to-G transition. In D-stage, CEN2-GFP signal seems normal. We now mention this in the text and in the Fig. 4 Supplementary 1B.

      (3) The authors nicely show the importance of MTs in the structure of the nest from which procentrioles and DEUP1 positive structures emerge. They suggest this nest may be what supports procentriole generation in the absence of DEUP1 and parental centrioles. Firstly how does this nest look in the absence of DEUP1 and/or parental centrioles (centrinone treatment)? This may be what they are trying to show in Figure 5 Supplement 1 but it currently is very difficult to digest what it is showing relative to controls and whether this is significant in the way it is plotted.

      The nest is conserved in the DEUP1KO with or without centrosomal centrioles, as shown by accumulation of Centrin and PCNT at the center of the self-organised MT network (Mercey et al., 2019). This is in fact what motivated our study on the role of MT in centriole amplification. We have edited the legend to precise the quantification done, which is not related to this question. In this quantification, we show that the increased propensity to accumulate PCNT by centriole-loaded deuterosomes between A and G-stage is maintained in the absence of deuterosomes, indicating that centrioles themselves accumulate/recruit PCNT.

      (4) Can you do CLEM on DEUP1-Ruby and these early foci at the cloud stage to see if they are visible at the ultrastructural level, relative to procentrioles, microtubules, and other electron-dense structures?

      We thank the reviewer for this question. We have done CLEM on the pericentrosomal cloud during very early steps of centriole amplification. This showed that DEUP1 early accumulation at the centrosome corresponds to a region rich in fibro granular aggregates, suggesting that DEUP1 may be translated here, through locally concentrated centriolar sattelites, known to be involved in local translation. Then, small deuterosomes and immature centrioles are formed, within this cloud of sattelites, confirming that the pericentrosomal cloud is a nest for centriole biogenesis (Fig. 2C-D + Fig. 2 Supplementary 2-6 for control and Fig. 4 Supplementary 3-4 for nocodazole treated cells). This also shows that immature deuterosomes are not necessarily round shaped, and can be deprived of centriole loading.

      (5) Check the scale bars- see Fig 4E. Check throughout.

      Done.

      (6) Figure 3 Supplement 1 and 2 don't match the legend and are likely reversed - which one is right?

      Done.

      (7) Technical issue - I couldn't play videos 6 or 16? Check these work.

      Done.

      (8) Nomenclature mammalian proteins- mouse or human- should be all caps DEUP1, PLK4, SAS6,etc. Watch your units- space between number and unit.

      This has been done.

      (9) Many of the graphs involve three biological replicates but why not plot the mean of each of the three experiments and do stats? The number of events measured may conflate the significance. Try using Superplots.

      Here is how we proceed: we count the number of occurrence of the phenotype we monitor, and the total number of cells. We apply a X<sup>2</sup> to test whether there is a significative difference between our replicates in each condition. If not, we pool the number of occurrence of the phenotype we monitor and the total number of cells for the 3 replicates, and for each condition. Finally we apply a X<sup>2</sup> between the different conditions. This is how we usually proceed to avoid comparing a mean of percentages. This is now explained in the methods.

      Minor points:

      (1) "DEUP1 is a centrosomal protein and assembles deuterosomes in the pericentrosomal region in brain MCC". I am not sure you have evidence that DEUP1 is a centrosomal protein. You don't seem to study the relationship between centrosomes and DEUP1? Rewrite this title and tone down this claim.

      This has been modified.

      (2) Why the crossbow micropattern (versus some other shape) - seems very specific but not discussed?

      We wanted a shape where centrosome is not localized at the center of mass of the nucleus. Among the corresponding patterns, the crossbow was the one where differentiating cells had less propensity to detach.

      (3) Figure 2 - are the foci of DEUP1 at the cloud stage smaller than at A stage? How do they grow? Measure the diameter at cloud stage, just after they leave the cloud and then once they move away from centrosomal cloud and each other. If so, and they do indeed grow in size from the cloud stage to the growth stage which I think your images suggest - do you envision this happening with the gradual addition of DEUP1 rather than fusion?

      Early deuterosomes are not easy to detect by light microscopy, because of accumulation of DEUP1 in the cloud. We did CLEM on the cloud of early A-stage cells to resolve the earliest deuterosomes which are often very small (see Fig. 2D, Fig. 2 Supplementary 2-6) suggesting that they grow, either by fusion, which we never observe in our movies at later A-stage, or by accretion of DEUP1. However, by light microscopy, we can detect very early but big deuterosomes, which we see splitting later on into smaller ones. So, we cannot conclude on the mechanism that regulate deuterosome size. This is now discussed in the discussion of the manuscript.

      You say in the discussion:

      "Consistently, we never observed fusion events of DEUP1 condensates in our time-lapse experiments. More importantly, we did FRAP experiments on endogenously tagged mRuby-DEUP1 in cells at the different stages of centriole amplification, and did not find significant recovery, supporting that centrosomal DEUP1+ foci and deuterosomes are not liquid-like structures (Figure 8 Supplementary 2)." How do you prove there is no fusion of deuterosomes?

      It is always difficult to prove the absence of something, we agree! But we did tens of movies with high temporal resolution and never observed fusion events. But, as we say in the previous question, the very early deuterosomes can be very small and we do not distinguish them from the DEUP1+ cloud by live imaging. So at this stage, we cannot say. But later on, during A- or G-stage and when deuterosomes are outside the cloud to be easily observed, we very often observe deuterosomes bumping into each others and stay in close contact for minutes, but then moving away. This, for us, supports the lack of fusion properties. But the question remains open. We now explain this in the manuscript and have added an example in video 28.

      If they are getting bigger as I think your imaging suggests from cloud to growth stage, then how is this happening?

      MT depolymerisation and dynein inhibition leads to the formation of very small deuterosomes. Dynein inhibition can even lead to a block in the formation of new deuterosomes suggesting that DEUP1 concentration is a crucial parameter for condensation into deuterosomes. Deuterosome growth may happen through oligomerization of DEUP1 molecules allowed by their dyne-independent concentration. Sorokin in 1968 proposed that a supersaturation of deuterosome components may lead to their solid crystallization into deuterosomes. Deuterosome size can also be regulated by a more complex molecular cascade, involving post-translational modifications of DEUP1 or PCM, such as phosphorylations driven by the cell cycle machinery. This would be consistent with the fact that deuterosomes are very big in the absence of CCNO, a cyclin required for entering the MCC cell cycle variant. This will need further investigations.

      I'm not sure FRAP actually proves fusion doesn't happen.

      Agreed, this is not what we wanted to say, we clarified. The FRAP experiment just suggests that it is not liquid-like.

      It is technically difficult to laser ablate individual or only subsets of deuterosomes...

      This is what was done but anyway, FRAP does not firmly show that deuterosome compartments are not liquid-like as we now precise.

      (4) How do you fix your cells for expansion as you have no preservation of cytoplasmic microtubules? You are saying that there is a "nest" of MTs but beta tubulin ONLY stains the cilia and centriole - why is this? Tyrosinated tubulin on regular confocal shows strong cytoplasmic staining. See Figure 3.

      Cytoplasmic microtubules do not preserve well through the expansion process. We did try a few different fixations and pre-extraction methods but they come at a trade-off to preserving centrioles. i.e. we could either preserve cytoplasmic tubes or centrioles but not both with the same processing method.

      (5) "PCNT puncta partially overlap with centrin (Figure 3 Supplementary 2C). At this stage, PLK4, the master regulatory kinase, and SAS6, one of the first centriolar components are either absent or present as small foci within the cloud, often on the wall of the parent centrioles (Figure 3B-C)." some arrows to highlight this would be useful - difficult to see?

      We have tried to make arrows on what is now Fig. 3 Supplementary 1 G, but there is to many CENTRIN colocalizing with PCNT. We have enhanced the contrast of the merge to make it more visible.

      (6) Figure 3I legend - what are the arrows pointing at? Yellow and white on inserts? ". Around the same time as tubulin, centrin is also recruited to procentrioles (Figure 3I). This stage is probably the stage that we previously documented as A"

      However you see centrin at DEUP1 foci in D, and you don't show any eg. SAS6 or PLK4 positive DEUP1+ structures lacking centrin specifically, centrin seems to be present on all the procentrioles in Figure 3I. Did I miss it where you show centrin negative procentrioles in the cloud?

      Fig. 3I (now Fig. Supplementary 1J), yellow arrows are pointing at centrioles with non-acetylated MT while white arrows point at acetylated MT. This is now indicated in the legend.

      Regarding CENTRIN, it is present as a diffuse staining around the centrosome since the very beginning of amplification (now in Fig. 3 Supplementary 1A with different contrasts), in addition to compose the parental centrioles. This staining can therefore overlap with DEUP1 staining when DEUP1 appears (Fig. 3 Supplementary 1B, E) but not necessarily. In live we observe that CENTRIN and DEUP1 foci can move independently at early stages (Fig. 2 Supplementary 1B, video 2). This is later on, as shown now in Fig. 3 Supplementary 1J (previously Fig. 3I), that procentrioles are all strongly positive for CENTRIN.

      A new paper (Laporte et al., Cell 2024) recently showed that the recruitment of CENTRIN on duplicating procentrioles first occurs at the distal end, visible by a small dot, and then appears gradually at the level of the inner scaffold when procentriole reach 160nm, the stage where POC5 appears, which corresponds to the A-to-G transition in our MCC progenitors (Al Jord et al., 2014). One can therefore consider that the same is happening in our cells, and that, with the CENTRIN cloud, we have difficulties to detect the distal CENTRIN dot. We have changed the text to add this reference and discuss CENTRIN apparition in MCC procentrioles.

      (7) " The DEUP1 asymmetry previously described at the centrosomal daughter centriole (Al Jord etal., 2014) becomes visible in some cells during the cloud stage (Figure 3B, N; Figure 3 Supplementary 2B) and in a majority of cells" difficult to see - maybe enlarge and single channel from Figure 3F-H in the supplemental Figure 3 to emphasise this?

      We have either changed the pictures or the contrast to be more representative with the quantifications. This is visible in Fig. 3A, D, E, G; Fig3. Supplementary 1E and now using correlative light and EM in Fig. 2 Supplementary 2, 3, 4 and Fig. 4 Supplementary 3-4. One has to look carefully at the daughter centriole (marked “dc”). We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both strutures are frequently sticked to each others on tens of nanometers.

      (8) Do you have videos of DEUP1 oscillations with nocodazole to show a lack of oscillations?

      We have now added videos of DEUP1 oscillations under nocodazole and dynapyrazole treatments.

      (9) "In addition, co-staining of centrioles and nuclear pore proteins show a tight colocalization(Figure 6 Supplementary 1B)." I see the colocalisation in panel 1 but less obvious with panel 2 maybe have some more zoomed in panels and some quantification of the colocalization? Is it more striking at the G stage than the D stage?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. There is too many centrioles and NPC, they cannot do otherwise than colocalize… What we want to say is that there is a tight connexion with the nuclear envelope. This is why we present a single z, to show that centrioles and NPC are on the same z-plane. This is also clearly visible for centrioles that are loaded on deuterosomes that are around the nuclear membrane in the XY plane. We also added a video to show an entire z-stack of this kind of staining.

      (10) "Indeed, SAS6 normally disappears from procentrioles when centrioles are docked, just beforeciliation (Al Jord et al., 2014). This suggests that centrioles were able to degrade SAS6, a process also dependent on APC/C (Strnad et al., 2007), but failed to disengage from deuterosomes." Figure 6 Supplement 1E-F - are you sure it wasn't that Sas6 wasn't loaded correctly at the earlier stage and so is reduced recruitment rather than premature disengagement of Sas6? If it is indeed premature disengagement of Sas-6 - what about CP110 - does the CP110 get loaded and is it still present in noco treated cells arrested in the D phase?

      We do not observe SAS6-negative procentrioles on deuterosomes at G-stage but only on deuterosomes in D-stage cells (cells with partly disengaged procentrioles). This is why we hypothesize that, because of the long duration of D-stage and knowing that SAS6 is finally degraded at the end of amplification (Al Jord et al., 2014), we are in the presence of cells where SAS6 has been degraded but where centrioles did not manage to disengage. This is now clarified in the text.

      (11) Can you track deuterostome splitting live? Maybe not enough spatial or time resolution?

      One has to monitor in 3D (multiple z because deuterosomes move a lot), 2 colors, high temporal resolution (dt=2-5’; to be able to track a single deuterosome), and long duration (deuterosomes are sometimes touching each other and then moving away, giving the impression that they split). This eventually leads to the bleaching of the mRuby fusion protein… We have put an example of what we think is a deuterosome splitting in Fig. 6E (former Fig. 7D). But we decided to finally monitor with low temporal resolution (dt=40’) to avoid photobleaching, and analyze numerous deuterosomes and cells to quantify the number and size of deuterosomes over time in single cells.

      (12) The MT nodes - can you segment the tyrosinated MTs and define nodes and then quantify theDEUP1 presence on them?

      Please see answer to reviewer 1 regarding this point.

      (13) Figure 8 supp 1 (E): Representative XY distribution of CEN2-GFP+ centrioles at the end of migration (Sas6 negative) in brain MCCs treated with DMSO, Nocodazole 1µM and 5µM (48h). Scale bar, 5µm Bit more detail on how you define fully migrated vs still migrating centrioles in z. You say you are using Sas-6 negativity to define fully migrated cells in the legend, yet you say noco treatment leads to premature sas-6 negativity, and yet the apical migration takes longer upon noco treatment?

      Nocodazole does not lead to premature SAS6 negativity but to a partial disengagement which lead to SAS6 negative “mature” centrioles being still connected to deuterosomes. We define complete migration when all the centrioles are on the apical side of the nucleus. We now clearly define what “apical” migration stands for in the main text and changed the pictograms in Fig. 8G to clarify this.

      (14) Figure 8H and video 18 - it isn't obviously clear to me that the noco-treated cells are "more erratic" or how you decide what counts as apically migrated successfully. How do you control for drift in z? Can you track individual centrioles as you did in untreated and define what is "erratic about their movement?

      Erratic means that the centrioles are moving away from each others, and back, in a non-predictable way, instead of migrating up and gathering. The drift in z of the whole cell is visible because there is always some centrioles, that are apically located at the beginning, that remains on the apical membrane, probably because they are already docked.

      We have indeed followed the centrioles individually in the nocodazole condition. However, in the control, the XYZ coordinates of one of the centrioles of the centrosome, which normally don’t move, are substracted to the coordinates of all the other centrioles as explained in the method section. This allows to have a subcellular reference, and to circumvent the movements of the cell, which are non-negligible at all at this timescale. In the nocodazole treated cells, the centrosomal centrioles share the erratic movements of the other centrioles and can migrate up and down, which exclude them as a reference. Since the nucleus is also moving a lot, we were left with no reference point.

      (15) Figure 8 supplement 1E can you quantify the final area of centriole patch in XY upon noco treatment?

      It was in main Fig. 8J and is now in Fig. 8 Supplementary 1F.

      (16) Figure 8J legend- MBB is never defined as an acronym.

      Thank you for pointing this.

      (17) Define what is the frequency and how is it calculated - Figure 8J.

      This is the MBB patch area in µm<sup>2</sup>

      Text edits:

      (1) "Altogether, these results suggest that, in this non-tissue-specific proxy of MCC progenitors, microtubules organize the onset of centriole amplification in the pericentrosomal region."

      Sentences have changed.

      (2) "Increasing the temporal resolution to 5-15s reveals that DEUP1+ foci observe an exhibit oscillatory dynamics to at the centrosome (Figure 2E, colored arrows, Video 3, 5/10 cells observed for 1-4min)."

      Sentences have changed.

      (3) "stage procentrioles were involved in this perinuclear migration and distribution. In fact, this dynamic is reminiscent of the centrosome migration that occurs during the G2-to-M progression in cycling cells in preparation for mitotic spindle organization. In cycling cells, this" Grammar - maybe change to "stage procentrioles were involved in this perinuclear migration and distribution. This is reminiscent of the centrosome migration that occurs during the G2-to-M".

      Sentences have changed.

      (4) "We then wondered whether these microtubule-dependent dynamics was were required for an efficient subsequent centriole disengagement during the following D-stage."

      Sentences have changed.

      (5) "Then, monitoring tens of disengagement movies, we identified a transient stage during which disengaging procentrioles redistribute isotropically in the 3 dimensions, along the nuclear membrane (Figure 6A, 4:30, Video 7) before losing its contact to migrate to the apical surface (Figure 6A, 6:30 to 14:00)."

      Sentences have changed.

      (6) Discussion: "Since pioneer electron microscopy studies on basal body production in quail oviduct MCC 35 years ago (Boisvieux-Ulrich et al., 1987, 1990; Boisvieux-Ulrich et al., 1989), this work is the first to assess the role of microtubules in the now finely described centriole amplification process. This"

      Sentences have changed.

      (7) "Using live imaging on brain MCC, we highlight the existence of a nest composed of DEUP1, PCNT and Centrin2, pre-assembled before the onset of centriole amplification onset."

      Sentences have changed.

      (8) "Recently, formation of DEUP1 pure condensates in solution as well as FRAP experiments after overexpression of DEUP1 in MCC progenitors suggested that deuterosomes where are not liquidlike structures (Yamamoto & Kitagawa, 2019). Consistently, we never observed fusion events of DEUP1."

      Sentences have changed.

      (9) "This reminds is reminiscent of the centriole-to-centrosome conversion occurring at the G2-M transition followed by the associated microtubule dependent nuclear migration of new centrosomes at mitosis onset (Agircan et al., 2014)."

      Sentences have changed.

      (10) "Following individual trajectories requires high resolutive resolution spatio-temporal live imaging while avoiding excessive light exposure which disturbs centriole migration (Boudjema et al., 2024)."

      Sentences have changed.

      (11) "Using high temporal resolution microscopy, we further identify that individual dynamics is are complex and can be splitted between divided into the baso-apical migration, where centrioles move in a processive and more..."

      Sentences have changed.

      Reviewer #3 (Recommendations For The Authors):

      (1) Growing MEF-MCCs on micropatterns has successfully mimicked the dynamics of centriole amplification in brain MCCs, allowing the authors to study the spatial origin of procentrioles. Since this is a powerful system, a more quantitative description of the system will be informative and beneficial for future studies. For example: What is the efficiency of this system? Do the cilia that form in MEF-MCCs motile?

      The system of MEF-MCCs has been described in a previous paper from the Kintner lab. It seems that growing the MEF-MCCs on micropatterns did not ameliorate the ciliation which is partial, probably due to the absence of an apico-basal polarity.

      (2) Figure 2: The analogy drawn by the authors between DEUP1 oscillatory dynamics and centriolar satellites is intriguing. In early amplifying cells within the cloud, do these DEUP1 structures co-localize with the satellite marker PCM1?

      We have added immuno stainings of PCM1 in mRuby-DEUP1 / CEN2-GFP cells in Fig. Supplementary 2E. Within the centrosomal cloud, DEUP1 colocalizes with PCM1. Interestingly, this PCM1 concentration at the centrosome is dependent, at least in part, on dyneins. Then, PCM1 can localize around the deuterosomes, but it is never colocalized with deuterosomes (not shown). This is also showed by immuno-EM in Zhao et al., 2019. Although it was shown that PCM1 is a proximity interactor of DEUP1 (called ccdc67 at that time) by Firat-Karalar et al., 2014., absence of PCM1 staining on deuterosomes does not favor the hypothesis of PCM1 and DEUP1 being part of the same entities. One could hypothesizes that DEUP1 is transcribed locally within the satellites, explaining the colocalization of the 2 proteins and the + BioID results, and then form PCM1negative deuterosomes.

      (3) The authors propose a physical link between deuterosomes and centrosomes based on their oscillatory behavior. How are the oscillatory dynamics of DEUP1 affected by nocodazole treatment or inhibition of microtubule motors (i.e ciliobrevin treatment)?

      These oscillations are inhibited by nocodazole (Fig. 4D). They are also inhibited by dynapyrazole (Fig. 4D). We never succeeded in having a nice disruption of the Golgi apparatus with ciliobrevin and therefore we did not used it.

      (4) In addition to nocodazole treatment, it would be important to determine the consequences of microtubule stabilization by taxol and inhibition of microtubule motors during critical stages of centriole amplification where microtubules are reported to play a role for the first time in this manuscript. Another interesting area of investigation will be to study the extent to which microtubule PTMs contribute to these processes.

      We now blocks dyneins during the different stages of amplification. The results are in main and associated Fig. 4, 5, 7, 8. The role of microtubule PTM, is not in the scope of this manuscript.

      (5) Describing microtubule dynamics along with Centrin/DEUP1 dynamics will be informative in assessing whether these structures associate and/or move along microtubules? Have the authors performed their imaging experiments with SIR tubulin?

      Yes, we have tried hard! But we have encountered different obstacles:

      3-color video microscopy is phototoxic,

      siRTubulin is bleaching very rapidly

      The density of microtubules in MCC makes the observation hardly informative

      (6) Figure 5: The role of PLK1 in centriole-centrosome conversion and generation of multiple MTOCs can be tested with a PLK1 inhibitor for further confirmation.

      We have also tried but inhibiting Plk1 blocks the A-to-G and G-to-D transitions so it was not possible to uncouple the role of Plk1 in stage transitions versus centriole maturation.

      (7) Figure 6: The tight co-localization of nuclear pore proteins with centrioles poses questions about the role of nuclear pore proteins or other nuclear proteins that are associated with centrioles during centriole disengagement and migration. Considering the existing literature on centrosome-nucleus attachments, can there be a way to test this question within the scope of this manuscript?

      We have tried to deplete Nup133 but it’s killing the cells. Our additional experiments now show that the nuclear migration of centrioles during G-stage is dynein dependent, reinforcing the parallel with centrosome migration in prophase. We also added results from our scRNA sequencing (Fig. 5 Supplementary 1) showing that some key players of centriole migration to the nuclear membrane are conserved in the MCC cell cycle variant, and expressed with a comparable dynamics as to the canonical cell cycle.

      (8) Figure 8: Manually tracking a subset of migrating centrioles to define their dynamics during centriole migration and docking provides valuable analysis for determining the molecular mechanism of these processes. In addition to microtubules, does actin contribute to this process? Since centrioles eventually migrate to the apical side in nocodazole-treated cells, there should be other molecular players involved in this process.

      We did block actin polymerization but we found that the different stages were affected and that it would be better to dedicate a whole manuscript on the role of actin during each stage of amplification. We discuss the migration mechanism, and the putative role of actin, in the discussion.

      (9) The legends for Supplementary Figures 1 and 2 in Figure 3 are mixed and need correction.

      Figures have been remodelled.

      (10) In Figure 3P, the term "PLK4+" is labeled in bright green, which is not clearly visible. It maybe beneficial to change the color of this label for better visibility.

      We have tried to correct this.

      (11) Figure 6F quantifies "% tethered flowers" on the nuclear membrane. When quantifying, is the3D localization of DEUP1 flowers in both DMSO- and Noc-treated cells considered? A flower may appear to be on the nucleus in 2D, but it could be detached from the membrane in a 3D view.

      The quantifications are done in 3D. However, flowers that are below or above the nucleus are not quantified since the space is confined and the resolution in z to small to see whether they are connected or not. This is now precised in the legend.

      Before the editors proceed with an updated assessment, they've requested that we pass on some of the comments that have arisen as part of the evaluation of your revised manuscript. They feel that these concerns should be addressed before we proceed with issuing a formal assessment and publishing the revised Reviewed Preprint:

      We thank the reviewers and the editors for the corrections and insighfull comments. We apologize for our delayed answer and hope our corrections in the main text and some of the figures will give them satisfaction.

      The revised manuscript is greatly improved with nice new data regarding the role of microtubules. It also has changed quite a bit including the title. The new focus is on the cell and centriole cycle variants in MCC. While this helped to focus the study, there remains an important issue related to the interpretation of the data and the proposed 2-in-1 cycle model. Before providing the final updated assessment, we ask you to address the following points (which were raised already in the first round of review): The manuscript still contains statements that are not aligned with published work and the current view in the field regarding the timing of events during canonical centriole biogenesis. These timings are in conflict with your model that 2 centriole cycles are "superposed" in the MCC cell cycle variant, as currently presented. An alternative straightforward interpretation would be that multiciliogenesis uses an accelerated centriole duplication cycle where key steps occur concomitantly or in short succession instead of being separated by mitotic divisions as in the canonical cycle.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we think we have proposed. When correcting our confusions as regard to centriole-to-centrosome conversion (as explained below) and putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (corrected Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a superposition of events that; although driven by the same molecular machinery, are normally occuring in two consecutive cell cycle. We explain ourself briefly in two paragraphs, before answering point by point to the questions of the reviewers.

      As regard to centriole-to-centrosome conversion:

      We thank the reviewer for pointing out that we used “MTOC conversion” for what is normally called “centrosome maturation”. We have removed the term “centriole-to-centrosome conversion” during the first round of revision but we now realize that “MTOC conversion” leads to the same misinterpretation as regard to the literature on centriole duplication.

      The reviewer asks us to refer to the work of the Tsou lab (Wang 2011, reference now added in the manuscript) showing that daughter centrioles are “modified” (e.g. recruit PCM, become competent for MT nucleation and duplication) during late M/early G1. This “centriole-to-centrosome conversion” can’t occur for our procentrioles at this stage since they are not even born during the mitosis that precedes MCC differentiation. Also, in our cells, such modification does not include the capacity to become competent for duplication since we know that procentrioles become basal bodies without making any round of duplication (Al Jord et al., 2014).

      Also, we have not done the experiments to tackle the question on when our centriole become “modified-like”. What we can say is that during A-stage, they become progressively positive for PCM (Fig. 5 Supplementary 2) and a weak signal shows that some MT are seen emerging from them (Fig. 5 and Fig. 5 Supplementary 2, and see point by point answer).

      What we do see is that, at the A-to-G transition, they increase their PCM recruitment, show clear and strong MTOC ability (sometimes as strong as the centrosomal centrioles), and that this is associated with migration and separation of centrosome/deuterosomes around the nuclear membrane (Fig. 5). We therefore connect this to what occurs at the G2/M transition which is an increased recruitment of PCM protein, an increased ability to nucleate MT, associated with centrosome migration and separation at the nuclear membrane. Since this process in the canonical cell cycle is called “centrosome maturation”, we therefore should refer to this term in our study. However, centrioles in the MCC variants are not organized in centrosomes, so we now compare what we see to the “centrosome maturation” of the canonical cell cycle with an associated reference (Joukov et al., 2018), but name it “centriole maturation”.

      We have modified the text (track changes visibles) and the schemes (Fig. 5, Fig. 5 Supplementary 1 and 2, Fig. 9, Fig. 9 Supplementary S1; new versions uploaded) accordingly.

      As regard to 1.5 or 2 cell cycles

      Except for the “MTOC conversion” that we have now changed, as explained above, we think our work does suggest (depicted on Fig. 9) what the reviewer states for centriole duplication: “In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs. Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium)”.

      We feel that going from early S to a G1 phase, after 2 mitosis, is what one can call “2 cell cycles”. One of the paper that inspired us a lot when studying how the cell cycle machinery can drive centriole amplification in MCC is a paper from Jadranka Loncarek team (Kong et al., 2014) where they also state that “nascent centrioles gradually mature through 2 cell cycles”. Very interestingly, in this study they show that when they enhance Plk1 activation, they could erase centriole age and new procentrioles are able to recruit PCM and appendages within only 1 cell cycle, without mitotic progression, like what we see in MCC. We have added the reference in our discussion.

      Point by point answer

      (1) Original work on canonical centriole disengagement and centriole-to-centrosome conversion should be cited (e.g. PMID: 16862117, PMID: 21576395)

      As explained earlier, we used the wrong term since the begining. We do not speak about the centriole-to-centrosome (nor MTOC) conversion since we do not test when centriole modification (Wang et al., 2011) occurs in the MCC cell cycle variant. We know that PCNT is present on the procentrioles during A-stage (as shown in Fig. 5 Supplementary 2B), but we do not know when it is recruited (UExM did not work properly with this antibody). We quantify a weak MT staining in regrowth experiment during A-stage and see that procentrioles can be connected to MT in both brain MCC and MEFs (as shown in Fig. 5D, E for brain MCC and Fig. 5 Supplementary 2F for MEFs) , but we do not know when during A-stage they become competent for nucleation. We therefore did not speak about this process that we do not document. What we clearly document/quantify is the enhanced MT nucleation capacities at the A-to-G transition, concomitent with the nuclear migration (easily defined with Cen2-GFP or GT335 stainings) and that we compare to centrosome maturation occuring at the canonical G2/M transition.

      (2) The authors state in several places that canonical centriole formation and maturation takes two iterations of the canonical cell cycle. This is imprecise. Based on the above work and work by others, the broadly accepted view is that it takes 1.5 cell cycles. This difference matters for the final proposed model (see below). Reviewed e.g. here: PMID: 20869612; PMID: 30601682

      Our answer is in the preamble.

      (3) "Centriole maturation cycle superposes with centriole elongation cycle in the MCC cell cycle variant": Your description of the canonical cycle differs from the current view in the field. In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). All this occurs in 0.5 cycles. Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs (total of 1.5 cell cycles). Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium).

      (4) Fig 5A, B and Fig. 9

      (a) Are 2 separate figures needed for the model? They seem redundant.

      We find it easier not to wait Fig. 9 to have the first part depicted.

      (b) The model shows loss of SAS6 throughout G1, but this already occurs during M/early G1

      Thanks. It was already ok in Fig. 9, we have modified for Fig. 5.

      The model shows "MTOC capacity/conversion" during S phase, but this occurs during early G1

      Thanks a lot, as explained earlier, we used the term MTOC conversion occurring in G1 for what is normally called centrosome maturation occurring in G2/M, as explained earlier. We do not speak anymore of MTOC conversion since we have not tackled this question (explained above). We have therefore removed MTOC conversion in the texts and the schemes and replaced it by “centrosome maturation” for the duplication cycle, and by “enhanced MT nucleation capacity” for the MCC cycle. To be clearer and schematize that procentrioles are competent for MT nucleation before G2/M or A/G transitions, we have added some MT nucleated from G1 procentrioles during the canonical cycle, and from late A-stage procentrioles during the MCC cycle.

      The model shows disengagement only in the second M phase, but this occurs already at the first M phase, directly following centriole biogenesis, right before centosome conversion.

      This is a big edition error in both Fig. 5 and 9. Of course the daughter centriole disengage during the first M-phase. This has been changed. Thanks a lot for spotting it. This, however does not contradict the hypothesis of superposition.

      We also added the acquisition of distal appendage which was written in Fig. 5 but not in Fig.

      9 for duplication during the second M-phase.

      When the correct timings are incorporated in the figure, the proposed superposition of two cycles is not an accurate description of the events. Instead, your data seem consistent with a model where MCC incorporates all steps in one cell cycle variant that lacks mitoses, so that disengagement and MTOC conversion occur together with centriole elongation, followed immediately by acquisition of DAs and SDAs.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we tried to propose. When putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a super opposition of events that; although driven by the same molecular machinery, are normally occurring in two consecutive cell cycle. This is notably consistent with the findings of Kong et al., 2014 cited previously.

      (5) While all reviewers felt that there was no need to introduce the new term "nest", they leave it to the authors to keep it. However, the authors may want to consider that the term is still not introduced and explained properly, which may confuse readers. For example, while this section reads like an introduction to the term: "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC", the term is already used two times before without explanation. The first mentioning is at the beginning of the results section and is followed by citations, which gives the impression that these studies describe the nest, which is not the case.

      The first mention of “nest” is in the end of introduction resuming the findings of the paper where the term is in the following context: “we found that centriole amplification emerges in a pericentrosomal “nest” concentrating core centriole/deuterosome elements”. We looked at nest definition in the Collins Dictionnary : “a structure or other place where creatures, esp. birds, give birth or leave their eggs to develop”, we felt this was clear. We added quotation marks around the term nest.

      Then, the result section opens with this sentence: “The origin of amplified centrioles in MCC remains controversial. Some live imaging experiments and electron microscopy suggest that the centrosome could constitute a nest for centriole and deuterosome biogenesis (Al Jord et al., 2014; Kalnins et al., 1972; Mori et al., 2017), but others have proposed that procentriole-loaded deuterosomes emerge independently from the centrosome location, all over the cytoplasm (Nanjundappa et al., 2019; Sorokin, 1968; Zhao et al., 2013, 2019).”. Here, the term nest is again used as a place of birth for centrioles and deuterosomes which is what is actually proposed in these papers. First, Kalnins el al., in 1969 (we made an error on the reference date, this has been changed), resume in their abstract “This observation suggests that all of the clusters may form initially in close association with the diplosomal centrioles”. Then, not to mention Al Jord 2014 which comes from our lab, the title of Mori et al. is “Cytoplasmic E2f4 forms organizing centres for initiation of centriole amplification during multiciliogenesis”, and in the paper, they show that E2F4 accumulates at the centrosome. This is now also proposed by collaborators for MCIDAS (Lu et al., 2025). We feel that these references, which are often omitted, are appropriated at this location.

      Then we continue with: “To test whether microtubules drive the organization of a centrosomal nest from which procentrioles emerge”, which keeps the notion of the place of birth.

      Then the title "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC" arrives. In this section we first speak about a pericentriosomal cloud on which we zoom in using CLEM, to then conclude at the end of the section “Altogether live imaging mRuby-DEUP1/CEN2-GFP during early A-stage suggests that core deuterosome and centriole components are concentrated in a primordial cloud around the centrosome, which constitutes a nest where centrioles and deuterosomes concomitantly form before they move away from the centrosomal region (Fig. 2F)”.

      Finally, we begin the discussion section regarding the nest by: “We named this transitory compartment a “nest” since deuterosomes and procentrioles emerge specifically in this region and grow while moving away from it.”

      During the first revision, we tried to make it clearer. If this is still not the case after and the reviewer has another proposition of definitions/phrasing, we will be glad to consider it.

      As replied to the other reviewer, the term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the transitory region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      The following comments from Reviewer #3 may also provide further context regarding the editors' remaining concerns:

      The authors have done an excellent job addressing the points I raised overall, and the revision is substantially improved in focus and clarity. That said, some concerns raised by other reviewers, particularly regarding terminology and statistical analysis, could have been addressed more fully. One issue remains insufficiently resolved. Several quantitative analyses (for example Fig. 5C and 5E) still appear to rely on pooled single-event measurements collected across three independent experiments. This approach can overstate statistical significance. The authors indicate in their rebuttal that they use chi-square tests to compare proportions and to justify pooling across replicates. However, I am not convinced this addresses the issue for the intensity-based and single event distributions shown in the panels specified above. I recommend that these key analyses be represented with biological replicates shown explicitly (superplot-style, with replicates distinguished).

      Our reply was for the comparison of proportions and not the intensity-based and single event distributions shown in the panels Fig. 5C and Fig. 5E. We have now changed our plots to represent biological replicates explicitly (superplot-style, with replicates distinguished). As for the statistical analysis: we evaluated differences in marker intensity between A-stage and G-stage samples using a linear regression model, with stages as the main effect and replicate as a fixed covariate, to account for batch variation. Statistical significance was assessed using Type II ANOVA.

      Separately, I continue to feel that some newly introduced terminology (for example, the "nest") may not be necessary at this stage. It may be sufficient to describe these structures and focus on their spatiotemporal behavior, composition, and measurable features, rather than assigning new names. Having read the authors' response, I understand that they would like to retain this terminology, which is acceptable; however, it may not be readily adopted by the field.

      The term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      Minor correction (remove "in MCCs" part from the following sentence):

      In MCC, PCM1 depletion alters deuterosome formation and centriole production in brain and airway MCC (Hall et al., 2023; Zhao et al., 2021).

      Done

    1. eLife Assessment

      This manuscript provides valuable insight into how genome organization changes as cells progress through the cell cycle after mitotic exit, identifying two sharp genome remodeling events at G1-S and to a lesser extent, at S-G2 transitions. The conclusions are supported by solid, rigorous data, including sequencing and orthogonal imaging data. The use of sorted unsynchronized cells rather than cells treated with drugs is a particular strength.

    2. Reviewer #1 (Public review):

      This work convincingly shows that, rather than gradually "evolving" throughout interphase, global chromatin architecture undergoes unexpectedly sharp remodeling at G1-S (and to a lesser extent, S-G2) transitions. By applying "standard" Hi-C analyses on carefully sorted cells, the authors provide an excellent temporal view of how global chromatin architecture is changed throughout the cell cycle. They show a surprisingly abrupt increase in compartmentation strength (particularly interactions between the "active" A compartments) at G1-S transition, which is slightly weakened at S-G2 transition. Follow-up experiments show convincingly that the compartment "maturation" does not require the DNA synthesis accompanying S phase per se, but the authors have not identified the responsible factors (work for future publications). The possible biological ramifications of these architectural changes (setting up potential replication "factories", and/or facilitating transcription-replication conflict resolution, both more pertinent for the active A compartments, which are most affected) have been well discussed in the article, but still remain speculative at this stage.

      My major criticism of this article is aimed more at the state of the field in general, rather than this specific article, but it should be discussed to give a more balanced view: what actually is a chromatin compartment? Chromosomal tracing and live tracking experiments have shown that the majority of "structures" identified from Hi-C experiments are statistical phenomena, with even "strong" interactions only being infrequent and transient. A-B compartments are "built up" from multiple very low-frequency "interactions", so ascribing causal effects for genome functions is even tougher. As a result, I have very little confidence in the results of the authors' polymer simulations and their inferred "peninsula" A compartment structures without any other supporting experimental data.

      Comments on revised version.

      The authors have included orthogonal DNA FISH evidence to support their claims which greatly strengthens the manuscript. Their further precisions within the discussion have answered all of my previous concerns with the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript by Choubani et al presents a technically strong analysis of A/B compartment dynamics across interphase using cell-cycle-resolved Hi-C. By combining the elegant Fucci-based staging system with in situ Hi-C, the authors achieve unusually fine temporal resolution across G1, S, and G2, particularly within the short G1 phase of mESCs. The central finding that A/B compartment strength increases abruptly at the G1/S transition, stabilizes during S phase, and subsequently weakens toward G2 challenges the prevailing view that compartmentalization strengthens monotonically throughout interphase. The authors further propose that this "compartment maturation" is triggered by S-phase entry but occurs independently of active DNA synthesis, and that it involves a consolidation and large-scale reorganization of A-compartment domains.

      Strengths:

      Overall, this is a thoughtfully executed study that will be of broad interest to the 3D genome community. The data are of high quality, and the analyses are extensive, albeit not completely novel. In particular, previous work (Nagano et al 2017 and Zhang et al 2019) has shown that compartments are re-established after mitosis and strengthened during early interphase, and single-cell Hi-C studies have reported changes in compartment association across S phase. In particular, Nagano et al show that DNA replication correlates with a build-up of compartments, similar to what is presented here, with the authors' conclusion that compartment strength peaks in early S. The idea that it weakens toward G2, rather than continuing to strengthen, appears to be novel and differs from the prevailing framing in the literature.

      Comments on revised version.

      The authors have responded constructively to my major conceptual concerns. The distinction between DNA synthesis and replication initiation has been clarified appropriately. The additional insulation analysis substantially strengthens the argument that compartment maturation is not simply a consequence of changing loop extrusion dynamics, although I would encourage slightly more cautious wording regarding "independence" from cohesin-mediated extrusion. The peninsula model is now framed appropriately as a heuristic interpretation and supported by orthogonal imaging data. Finally, the discussion of conservation across cell types has been appropriately tempered. Overall, I believe the manuscript has been significantly improved.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work convincingly shows that, rather than gradually "evolving" throughout interphase, global chromatin architecture undergoes unexpectedly sharp remodeling at G1-S (and to a lesser extent, S-G2) transitions. By applying "standard" Hi-C analyses on carefully sorted cells, the authors provide an excellent temporal view of how global chromatin architecture is changed throughout the cell cycle. They show a surprisingly abrupt increase in compartmentation strength (particularly interactions between the "active" A compartments) at G1-S transition, which is slightly weakened at S-G2 transition. Follow-up experiments show convincingly that the compartment "maturation" does not require the DNA synthesis accompanying S phase per se, but the authors have not identified the responsible factors (work for future publications). The possible biological ramifications of these architectural changes (setting up potential replication "factories", and/or facilitating transcription-replication conflict resolution, both more pertinent for the active A compartments, which are most affected) have been well discussed in the article, but still remain speculative at this stage.

      We thank Reviewer #1 for their positive and constructive assessment of our work, and we agree that the questions of responsible factors and biological ramifications are important directions for future studies.

      My major criticism of this article is aimed more at the state of the field in general, rather than this specific article, but it should be discussed to give a more balanced view: what actually is a chromatin compartment? Chromosomal tracing and live tracking experiments have shown that the majority of "structures" identified from Hi-C experiments are statistical phenomena, with even "strong" interactions only being infrequent and transient. A-B compartments are "built up" from multiple very low-frequency "interactions", so ascribing causal effects for genome functions is even tougher. As a result, I have very little confidence in the results of the authors' polymer simulations and their inferred "peninsula" A compartment structures without any other supporting experimental data.

      We thank the reviewer for raising this important conceptual point. This issue extends beyond the scope of the present study but reflects an important ongoing discussion in the 3D genome field regarding the biological interpretation of chromatin compartments.

      We agree that Hi-C interactions should not be interpreted as stable pairwise contacts present in every cell. A growing body of evidence from chromatin tracing and live-cell imaging studies has demonstrated that many chromatin interactions identified by Hi-C are probabilistic and dynamic, with substantial cell-to-cell variability. Relatively speaking, however, A/B compartment organization represents a robust population-level property of genome organization that is highly reproducible across biological replicates and closely correlates with multiple independent genomic features. In particular, replication timing (RT) correlates very well with A/B compartment organization, with early and late RT domains corresponding to A and B compartment domains, respectively.

      Furthermore, single-cell DNA replication sequencing (scRepli-seq) analyses have revealed remarkably low cell-to-cell variability in RT, suggesting that RT profiles and A/B compartment organization reflect biologically meaningful and relatively stable features of nuclear architecture rather than purely statistical artifacts. Thus, while individual chromatin contacts may be transient and probabilistic, the megabase-scale compartment organization inferred from them appears sufficiently reproducible to support reproducible RT programs and other genome functions. Additional support comes from decades of work on DNA replication demonstrating that spatiotemporal replication patterns, visualized as replication foci following short EdU pulses, are remarkably reproducible between individual cells throughout S-phase progression. These patterns reveal clear spatial segregation between early-replicating A-compartment regions and late-replicating B-compartment regions even at the single-cell level.

      To directly address the reviewer’s concern that A/B compartment organization might represent only an ensemble-level statistical phenomenon without biological relevance at the single-cell level, we performed L1/B1-EdU DNA FISH on asynchronous mESCs and MC12 embryonic carcinoma cells. L1 elements are enriched in B compartment domains, while B1 elements are enriched in A compartment domains, allowing visualization of compartment segregation in individual nuclei across the cell cycle. This single-cell analysis confirmed our Hi-C findings: compartment segregation increased from G1 to early S, remained elevated throughout S phase with reduced cell-to-cell variability, and then weakened in G2. Thus, compartment segregation is detectable in single cells, and the temporal dynamics of compartment maturation identified by population Hi-C were independently recapitulated at single-cell resolution. We have added a new Results section describing these findings titled “Stepwise A/B compartment reorganization during interphase is conserved at single-cell resolution”, including new Figure panels 2D–H and Figure S5.

      Regarding the polymer simulations, we agree that these models should be interpreted with caution. We do not view them as direct representations of individual nuclei, but rather as heuristic models that help visualize structural trends present in the Hi-C data. To make this point explicit, we have added the following statement to the revised manuscript: “We note that these models are derived from population-averaged Hi-C data and should therefore be interpreted as a heuristic framework for understanding A/B compartment dynamics, rather than as definitive representations of individual nuclei.”

      That said, we did try to provide orthogonal experimental support for the "A peninsula" model by performing DNA FISH. In brief, we measured distances between probe pairs spanning two A domains on chromosomes 2 and 15 across different cell-cycle stages. We observed significant increases in inter-probe distances from G1 to early/mid S, with the most pronounced changes involving the central probes (i.e., probes located near the domain center), consistent with physical extension of the A domain during S phase. While these data do not prove the exact geometry depicted by the model, these findings provide independent experimental support for the peninsula model as a simplified but biologically grounded interpretation of the Hi-C data. These results are described in the Results section titled “A-compartment consolidation during S-phase involves enhanced long-range contacts and structural reorganization” and are presented in new Figure panels 5D–F and Figure S12.

      We thank the reviewer again for raising this important conceptual issue, which prompted us to better clarify both the biological interpretation and the limitations of our analyses.

      Specific minor points:

      (1) A better explanation for how Figure 1E was generated is required, because this figure could be very misleading. Figure 1F and all other cis-decay plots (and the Hi-C maps themselves) show that the strongest interactions are always at smaller genomic separations, so why should there be more "heat" at the megabase ranges in Figure 1E?

      We appreciate the reviewer's observation. The apparent discrepancy is simply due to the fact that the decay plot (Fig. 1E in the original submission, now Fig. S2C) does not include the shortest-range interactions. The lowest distance plotted is 25 kb, following the method originally described in Nagano et al. (Nature, 2017), which we used as a reference. The shortest-range interactions (below 25 kb) are indeed the most enriched, as seen on the diagonal of the Hi-C maps (Fig. 2A) and in the standard cis-decay plot (Fig. 1F in the original submission, now Fig. S2F). With the 25 kb cutoff in place, the "heat" observed at megabase distances (specifically 12–50 Mb) in early/mid G1 corresponds to the dark, non‑specific band around the diagonal visible in the Hi-C maps at the same time points. This is also reflected in the cis-decay plot (Fig. S2F), where distances in that range appear above the expected curve (a "bump" rather than a linear decay).

      To avoid confusion, we have updated the figure legend accordingly (Fig. S2C): “(C) Contact decay profiles for all cell cycle phases, plotted from 25 kb to 50 Mb, illustrating a continuum of cis-interactions and a progressive shift from long-range (> 12 Mb) to short-range (< 1 Mb) interactions during the G1-to-S phase transition.”

      We hope this explanation clarifies the figure.

      (2) An ultra-high-resolution Hi-C study (Harris et al., Nat Commun, 2023) identified very small A and B compartments, including distinctions between gene promoters and gene bodies, raising further questions as to what the nature of a compartment really is beyond a statistical phenomenon. It is unreasonable to expect the authors to generate maps as deep as this prior study, but how much do their conclusions change according to the resolution of their compartment calling? The authors should include a balanced discussion on the "meaning" of A/B compartments.

      We thank the reviewer for highlighting recent ultra-high-resolution work, such as Harris et al. (Nat Commun, 2023), which reveals compartment-like features at much finer genomic scales. We agree that these findings raise important questions regarding the scale-dependence and interpretation of A/B compartmentalization.

      In our study, we specifically focus on coarse-grained compartment organization, analyzed across multiple resolutions (from ~1 Mb to sub‑megabase scales). Importantly, the key conclusions, including the abrupt strengthening of compartmentalization at the G1/S transition, are robust across these resolutions.

      We also note that fine-scale compartment-like features likely operate under different rules than larger-scale compartments. Recent evidence suggests that these "micro‑compartments" are more dynamic and transient (Harris et al., Nat Commun, 2023; Goel et al., Nat Struct Mol Biol, 2025), whereas the large-scale compartments analyzed here capture more stable, global segregation patterns. Understanding how these two regimes relate to one another remains an important open question.

      We have added the following statement in the Discussion acknowledging the scale-dependent nature of compartmentalization: “At the same time, recent ultra-high-resolution Hi-C studies [36,37] have revealed compartment-like features at much finer genomic scales, emphasizing that A/B compartmentalization is, to some extent, inherently scale-dependent. Understanding how these fine-scale, often transient micro-compartments relate to the more stable, large-scale segregation patterns described here will be an important direction for future studies.”

      Reviewer #2 (Public review):

      Summary:

      This manuscript by Choubani et al presents a technically strong analysis of A/B compartment dynamics across interphase using cell-cycle-resolved Hi-C. By combining the elegant Fucci-based staging system with in situ Hi-C, the authors achieve unusually fine temporal resolution across G1, S, and G2, particularly within the short G1 phase of mESCs. The central finding that A/B compartment strength increases abruptly at the G1/S transition, stabilizes during S phase, and subsequently weakens toward G2 challenges the prevailing view that compartmentalization strengthens monotonically throughout interphase. The authors further propose that this "compartment maturation" is triggered by S-phase entry but occurs independently of active DNA synthesis, and that it involves a consolidation and large-scale reorganization of A-compartment domains.

      Strengths:

      Overall, this is a thoughtfully executed study that will be of broad interest to the 3D genome community. The data are of high quality, and the analyses are extensive, albeit not completely novel. In particular, previous work (Nagano et al 2017 and Zhang et al 2019) has shown that compartments are re-established after mitosis and strengthened during early interphase, and single-cell Hi-C studies have reported changes in compartment association across S phase. In particular, Nagano et al show that DNA replication correlates with a build-up of compartments, similar to what is presented here, with the authors' conclusion that compartment strength peaks in early S. The idea that it weakens toward G2, rather than continuing to strengthen, appears to be novel and differs from the prevailing framing in the literature.

      We thank Reviewer #2 for their thoughtful assessment and critique. We address their specific concerns below.

      Weaknesses:

      That said, several aspects of the conceptual framing and interpretation would also benefit from further clarification, and the mechanistic interpretation of the reported compartment dynamics requires more careful positioning relative to established models of genome organization. Specific concerns are outlined below:

      (1) One of the major conclusions of the study is that compartment maturation does not require ongoing DNA replication. However, the interpretation would benefit from more precise wording. Thymidine arrest still permits licensing, replisome assembly, and other S-phase-associated chromatin changes upstream of bulk DNA synthesis. Therefore, their data, as presented, demonstrate independence from DNA synthesis per se, but not necessarily from the broader replication program. Please clarify this distinction in the text and interpretations throughout the manuscript.

      We thank the reviewer for this important distinction. We agree with their point and have never claimed that compartment maturation is independent of the broader replication program. That is why we carefully used the term "active DNA synthesis" rather than "replication" throughout the manuscript.

      However, we acknowledge that one sentence in the text was ambiguous. The original sentence read: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a pre-replicative state where replication had not yet initiated, although cell-cycle markers indicated entry into S-phase.”

      We have now revised it to: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a state where the replication program (including origin licensing, replisome assembly, and helicase activation) has been initiated, as indicated by cell-cycle markers, but ongoing DNA synthesis (elongation) is blocked. ”

      This clarifies that compartment maturation is independent of active DNA synthesis (elongation) but not necessarily independent of upstream replication-associated processes. The change has been made in the manuscript.

      (2) A major conceptual issue that is not addressed at all is the well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartmentalization. Numerous studies have shown that loss of cohesin or reduced loop extrusion leads to stronger compartment signals, whereas increased cohesin residence or enhanced extrusion weakens compartmentalization. Given this framework, an obvious alternative explanation for the authors' observations is that the abrupt increase in compartment strength at G1/S, and its decline toward G2, could reflect cell-cycle-dependent modulation of cohesin activity rather than a compartment-intrinsic "maturation" program.

      The manuscript does not explicitly consider this possibility, nor does it examine loop extrusion-related features (such as loop strength, insulation, or stripe patterns) across the same cell-cycle stages. Without discussing or analyzing this widely accepted model, it is difficult to distinguish whether the reported compartment dynamics represent a novel architectural mechanism or an indirect consequence of known changes in extrusion behavior during the cell cycle. I strongly encourage the authors to analyze their data to determine if they observe anti-correlated loop changes at the same time they observe compartment changes. Ideally, the authors would remove loop extrusion during interphase using well-established cohesin degrons available in mESCs and determine if the relative differences in compartment dynamics persist.

      We thank the reviewer for raising this interesting point. We agree that there is a well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartment strength in the literature.

      To test whether cell cycle compartment dynamics, particularly compartment maturation at the G1/S transition, could be explained by changes in loop extrusion, we analyzed insulation at RAD21/CTCF sites (mESC data from Hansen et al., eLife, 2017) across the cell cycle. During normal cycling, we indeed observed an anti-correlation: insulation dropped as compartment strength increased at the G1/S transition. However, in G1/S-arrested cells, insulation did not drop compared to late G1 (it even slightly increased) even though compartment maturation still occurred, indicating that the two processes can be uncoupled. This is consistent with other studies showing that loop extrusion and compartment dynamics are driven by independent mechanisms (Nora et al., Cell, 2017; Zhang et al., Nat Commun, 2021), although we cannot fully rule out some contribution from loop extrusion dynamics without direct cohesin degron experiments.

      We have added a new Results section describing these findings titled “Compartment maturation is independent of cohesin-mediated loop extrusion”, including new Figure panels 3H, I, and Figure S7.

      (3) The proposed "peninsula-like" A-domain structures are inferred from ensemble Hi-C data and polymer modeling, rather than directly observed physical conformations. That is, single-cell imaging data clearly have shown that Hi-C (especially ensemble Hi-C) cannot uniquely specify physical conformations and that different underlying structures can produce similar contact patterns. The "peninsula" language, as written, risks being interpreted as a literal structural model rather than a conceptual visualization. Instead of risking this as just another nuanced Hi-C feature in the field, the authors could strengthen the manuscript by either (i) explicitly framing the peninsula model as a heuristic description of contact redistribution rather than a definitive physical architecture, or (ii) discussing alternative structural scenarios that could give rise to similar Hi-C patterns. Clarifying this distinction would improve the rigor and help readers better understand what aspects of A-compartment consolidation are directly supported by the data versus model-based extrapolations. For example, it would be useful to clarify whether the observed increase in long-range A-A contacts reflects spatial extension of internal A regions, changes in loop extrusion dynamics, increased compartment mixing within the A state, or population-averaged heterogeneity across alleles.

      We thank the reviewer for this important clarification. We agree that the "peninsula" model should be framed as a heuristic description. As detailed in our response to Reviewer #1 (see above), we have added a disclaimer to the manuscript and provided orthogonal DNA FISH support for physical extension of A-domains during S phase. We have also ensured that the language emphasizes the conceptual nature of the model.

      (4) The extension of the analysis to additional cell types using HiRES single-cell data is a valuable addition and supports the idea that compartment maturation is not unique to mESCs. However, the limitations of these data, in particular, the limited phase resolution, in addition to the pseudo-bulk aggregation and variable coverage, should be emphasized more clearly in the main text. Framing these results as evidence for conservation in principle, rather than definitive proof of identical dynamics across tissues, would be a more appropriate framing.

      We agree with the reviewer. We have already explicitly acknowledged the limited temporal resolution and variable coverage of the HiRES dataset in the main text. To better reflect its supporting role, we have moved the HiRES figure (previously Fig. 4) to Fig. S10 and merged the corresponding results section with the previous one titled: “Formation of a consolidated A compartment in S-phase”.

      We have also revised the language to avoid overstatement. The original conclusion read: “Together, these findings strongly indicate that compartment maturation and the accompanying A compartment consolidation represent a robust and universally observed feature across different developmental contexts.”

      This has been changed to: “Together, these findings support the notion that compartment maturation and the accompanying A-compartment consolidation are not unique to mESCs and may represent a broadly conserved feature of mammalian chromatin organization.”

      Similarly, the abstract has been adjusted from: “Moreover, compartment maturation was not limited to mESCs but was also observed across different developmental contexts in mice.” to: “Moreover, compartment maturation was not limited to mESCs but was also evident across different developmental contexts in mice.”

      These changes frame the results as evidence for conservation in principle rather than definitive proof of identical dynamics across tissues.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please address the minor points in the public review.

      In addition, on page 7, line 285: "In contrast, interactions showed minimal change across all distances though interphase". Do the authors mean "In contrast, B-B interactions..."?

      We thank the reviewer for catching this. The sentence has been corrected.

    1. eLife Assessment

      This important study identifies PRRT2 as an auxiliary regulator of Nav channel slow inactivation in vitro and in vivo, proposing that PRRT2 facilitates entry into, and delays recovery from, the slow-inactivated state. The revised manuscript has been substantially strengthened, providing compelling evidence that PRRT2 is relevant to normal brain physiology and disease pathophysiology, providing a mechanistic link between PRRT2 mutations and episodic neurological phenotypes. Overall, this study will be of interest to ion channel biophysicists and neurophysiologists, particularly those studying channelopathies.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Lu and colleagues demonstrate convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      Comments on revised version.

      The manuscript by Lu and colleagues has been revised sufficiently to address all my prior concerns.

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

    3. Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".<br /> PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last (20th) compound APs in panels B and C.

    4. Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      (4) The mechanistic separation between trafficking of PRRT2 and its gating effects is not clearly resolved.

      (5) Additional studies with Nav1.6 should be carried out.

      Comments on revised version.

      These comments have been addressed in the revised version.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Lu and colleagues demonstrates convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously, including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      We thank the reviewer for these positive comments and for the thoughtful evaluation of our work.

      Weaknesses:

      There are a few missing experiments and one place where data are over-interpreted.

      (1) An in vitro study of Nav1.6 is conspicuously absent. In addition to being a major brain Na channel, Nav1.6 is predominant in cerebellar Purkinje neurons, which the authors note lack PRRT2 expression. They speculate that the absence of PRRT2 in these neurons facilitates the high firing rate. This hypothesis would be strengthened if PRRT2 also enhanced slow inactivation of Nav1.6. If a stable Nav1.6 cell were not available, then simple transient co-transfection experiments would suffice.

      We thank the reviewer for raising this point. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms.

      We have now performed new heterologous expression experiments to test whether PRRT2 modulates Nav1.6 slow inactivation. Consistent with our findings for other Nav isoforms, PRRT2 significantly enhances the slow inactivation of Nav1.6. We have incorporated these data into the revised Results and Figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) To further demonstrate the physiological impact of enhanced slow inactivation, the authors should consider a simple experiment in the stable cell line experiments (Figure 1) to test pulse frequency dependence of peak Na current. One would predict that PRRT2 expression will potentiate 'run down' of the channels, and this finding would be complementary to the biophysical data.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we performed a pulse-train protocol in the stable Nav1.2 cell line and quantified the use-dependent attenuation (“run-down”) of peak sodium current across successive depolarizations (Figure 1-figure supplement 1C). Compared with control cells, PRRT2-expressing cells exhibited a larger decline in peak current during trains, indicating greater reduction in channel availability during repetitive depolarizations (Figure 1-figure supplement 1C). This pattern is consistent with our observations above showing that PRRT2 enhances Nav channel slow inactivation. These new data have been incorporated into the revised manuscript. Please refer to Page 5, Lines 133-140; Figure 1-figure supplement 1C.

      (3) The study of one K channel is limited, and the conclusion from these experiments represents an over-interpretation. I suggest removing these data unless many more K channels (ideally with measurable proxies for slow inactivation) were tested. These data do not contribute much to the story.

      We agree with the reviewer’s assessment. To avoid over-interpretation and to maintain focus on PRRT2-dependent regulation of Nav channel slow inactivation, we have removed the potassium channel dataset and the associated conclusions from the revised manuscript.

      (4) In Figure 2, the authors should confirm that protein is indeed expressed in cells expressing each truncated PRRT2 construct. Absent expression should be ruled out as an explanation for the enhancement of slow inactivation.

      We thank the reviewer’s concern regarding expression of the truncated PRRT2 constructs in the Nav1.2 stable cell line, particularly PRRT2(1-266), which shows little effect on slow inactivation of Nav1.2 channels. In the revised manuscript, we conducted western blot to verify expression of the PRRT2(1-266)-HA construct in the Nav1.2 stable cell line. We have added these results to the revised manuscript, please refer to Page 6, Lines 171-173; Figure 2-figure supplement 1A and B.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is primarily expressed in the nervous system and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 directly interacts with Nav1.2 and Nav1.6, modulating channel properties and neuronal excitability.

      In this study, Lu et al. reported that PRRT2 is a physiological regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect can be replicated by the C-terminal region (256-346) of PRRT2, and is highly conserved across species from zebrafish, mouse, to human PRRT2. TRARG1 and TMEM233, the other two DspB family members, showed similar effects on Nav1.2 slow inactivation. Co-IP data confirms the interaction between Nav channels and PRRT2. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared to WT mice.

      Strengths:

      (1) This study is well designed, and data support the conclusion that PRRT2 is a potent regulator of slow inactivation of Nav channels.

      (2) This study reveals similar effects on Nav1.2 slow inactivation by PRRT2, TMEM233, and TRARG1, indicating a common regulation of Nav channels by DspB family members (Supplemental Figure 2). A recent study has shown that TMEM233 is essential for ExTxA (a plant toxin)-mediated inhibition on fast inactivation of Nav channels; and PRRT2 and TRARG1 could replicate this effect (Jami S, et al. Nat Commun 2023). It is possible that all three DspB members regulate Nav channel properties through the same mechanism, and exploring molecules that target PRRT2/TRARG1/TMEM233 might be a novel strategy for developing new treatments of DspB-related neurological diseases.

      We thank the reviewer for careful evaluation and insightful suggestions.

      Weaknesses:

      (1) Previously, the authors have reported that PRRT2 reduces Nav1.2 current density and alters biophysical properties of both Nav1.2 and Nav1.6 channels, including enhanced steady-state inactivation, slower recovery, and stronger use-dependent inhibition (Lu B, et al. Cell Rep 2021, Fig 3 & S5). All those changes are expected to alter neuronal excitability and should be discussed.

      We thank the reviewer for this suggestion. Although the present study focuses on PRRT2-dependent regulation of slow inactivation, we agree that PRRT2 may influence excitability through additional Nav-dependent mechanisms, including reduced current density and shifts in the voltage dependence of channel inactivation (Fruscione et al., 2018; Lu et al., 2021; Valente et al., 2023). Notably, because PRRT2 facilitates entry of Nav channels into slow-inactivated states both from closed states and from open states during prolonged depolarization, some of these previously reported effects may partly reflect enhanced slow inactivation and the resulting reduction in Nav channel availability. We have expanded the Discussion to integrate these prior findings and to clarify that these additional PRRT2-dependent effects may converge to shape neuronal excitability. Please refer to Page 16, Lines 445-452.

      (2) In this study, the fast inactivation kinetics was examined by a single stimulus at 0 mV, which may not be sufficient for the conclusion. Inactivation kinetics at more voltage potentials should be added.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we expanded our analysis of Nav1.2 fast-inactivation kinetics to include a range of test potentials (-20, -10, 0, +10, +20 and +30 mV) in the presence and absence of PRRT2. These experiments showed that PRRT2 expression did not significantly affect Nav1.2 fast-inactivation kinetics under these conditions. We have incorporated these new results into the revised manuscript. Please refer to Page 4, Lines 100-103; Figure 1C.

      (3) It is a little surprising that there is no difference in Nav1.2 current density in axon-blebs between WT and Prrt2-mutant mice (Figure 7B). PRRT2 significantly shifts steady-state slow inactivation curve to hyperpolarizing direction, at -70 mV, nearly 70% of Nav1.2 channels are inactivated by slow inactivation in cells expressing PRRT2 when compared to less than 10% in cells expressing GFP (Figure supplement 1B); with a holding potential of -70 mV, I would expect that most of Nav channels are inactivated in axon-blebs from WT mice but not in axon-blebs from Prrt2-mutant mice, and therefore sodium current density should be different in Figure 7B, which was not. Any explanation?

      We thank the reviewer for raising this point. In our axonal bleb recordings, although the holding potential was -70 mV, sodium current density was measured after a hyperpolarizing pre-pulse to -110 mV, which was applied before the test depolarization to relieve inactivation as much as possible (as described in the Methods). Therefore, the current density measurement in Figure 7B reflects the available current after this recovery step, rather than the steady-state availability at -70 mV. The lack of a difference in Figure 7B does not contradict the PRRT2-dependent shift in steady-state slow inactivation. In the revised manuscript, we have clarified this point explicitly in the Results and figure legend to avoid confusion. Please refer to Page 10, Lines 294-295.

      (4) Besides Nav channels, PRRT2 has been shown to act on Cav2.1 channels as well as molecules involved in neurotransmitter release, which may also contribute to abnormal neuronal activity in Prrt2-mutant mice. These should be mentioned when discussing PRRT2's role in neuronal resilience.

      We thank the reviewer for this suggestion. In addition to the Nav-dependent mechanisms, previous studies have shown that PRRT2 also regulates synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018) and presynaptic surface expression of Cav2.1 channels (Ferrante et al., 2021). These effects are also expected to influence neurotransmitter release and, consequently, neuronal and network excitability. In the revised manuscript, we have expanded the Discussion to acknowledge that these additional PRRT2-dependent mechanisms may also contribute to cortical resilience. Please refer to Page 16, Lines 452-457.

      Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      We thank the reviewer for this positive evaluation of our work and for the constructive comments.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav channel interface—through approaches such as targeted mutagenesis, crosslinking, and structural determination—will be required to define the binding interface and establish the molecular basis of gating modulation. Please refer to Page 16, Lines 465-468.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      We agree with the reviewer. Impaired slow inactivation in Prrt2-mutant mice is one plausible contributor to reduced cortical resilience. PRRT2 has also been reported to regulate surface exposure of Nav and Cav2.1 channels (Ferrante et al., 2021), as well as neuronal synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018). Each of these PRRT2-associated processes could influence cortical excitability in vivo. We have therefore expanded the Discussion to clarify that the cortical phenotype likely reflects the combined contribution of multiple PRRT2-dependent mechanisms, rather than an isolated defect in slow inactivation alone. Please refer to Page 16, Lines 446-458.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      We thank the review for this comment regarding physiological relevance. In the revised manuscript, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Please refer to Page 14 and 15, Lines 414-416; Lines 429-430.

      (4) The mechanistic separation between the trafficking effect of PRRT2 and its gating effects is not clearly resolved.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. Previous studies in heterologous overexpression systems have shown that PRRT2 can influence Nav channel trafficking and surface expression, raising the possibility that the observed effects on slow inactivation regulation might be secondary to altered channel abundance or localization. However, slow inactivation develops on a timescale of tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer intervals (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). These distinct temporal profiles argue against trafficking as the primary basis for the effects of PRRT2 on Nav channel slow inactivation described here, although direct quantification of dynamic changes in Nav channel surface expression will be required to fully exclude such a contribution (Liu et al., 2022; Tyagi et al., 2025). We have incorporated this point into the Discussion section. Please refer to Pages 13, Lines 378-388.

      (5) Additional studies with Nav1.6 should be carried out.

      We thank the reviewer for this suggestion. We have performed experiments to directly examine the effects of PRRT2 on Nav1.6 slow inactivation and incorporated these new data into the revised Results and figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for future experiments (not for this paper)

      (1) Exploit the lower protein expression in V5-PRRT2 mice to examine the effects of a hypomorphic allele.

      We thank the reviewer for this insightful suggestion. We note that the V5 epitope knock-in reduced PRRT2 protein expression, which may functionally resemble a hypomorphic allele. Accordingly, in addition to its utility for biochemical experiments (e.g., co-immunoprecipitation), this line could serve as a genetic tool to interrogate PRRT2 dose-dependent effects in vivo. We have added this point to the revised manuscript, please refer to Page 9, Lines 265-267.

      (2) Examine disease-causing PRRT2 mutations.

      We thank the reviewer for this constructive suggestion. Testing disease-associated PRRT2 variants for their ability to regulate Nav channel slow inactivation would be an important next step to strengthen the disease relevance of the mechanism proposed here. Moreover, identifying missense variants that selectively disrupt slow-inactivation regulation could help pinpoint residues that are critical for PRRT2-Nav functional coupling and thereby inform future structure-function studies. We plan to pursue this direction in follow-up work.

      (3) Investigate spreading depolarization in PRRT2-deficient mice.

      We thank the reviewer for this suggestion. Although we have shown that PRRT2 deficiency facilitates spreading depolarization in the cerebellum, whether PRRT2 exerts similar control over spreading depolarization susceptibility in the cerebral cortex remains to be determined. We plan to address this in an independent study and to test how cortical spreading depolarization relates to other PRRT2-associated neurological disorders.

      Reviewer #2 (Recommendations for the authors):

      This study is, in general, well executed, and the manuscript is well written. However, I do have some questions.

      (1) The authors' previous works have shown that PRRT2 regulates both Nav1.2 and Nav1.6, considering the wide expression Nav1.6 in CNS and its role in neuronal activity, what makes the authors not include Nav1.6 in this study?

      We thank the reviewer for raising this question. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms. In response to reviewers’ concern, we have now performed new experiments to directly examine the effect of PRRT2 on Nav1.6 slow inactivation. These results have been incorporated into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) Please explain why you chose 0 mV rather than -70 mV (closer to membrane potential) in the slow inactivation protocol.

      We thank the reviewer for raising this question. Nav channels can enter into slow inactivation from both resting/closed states and activated/open states. In our steady-state slow-inactivation assays, we found that PRRT2 enhances Nav1.2 slow inactivation under both conditions (Figure 1-figure supplement 1A and B). In whole-cell recordings, Nav1.2 channels typically begin to activate at command voltages more depolarized than approximately -60 mV. Accordingly, a conditioning voltage of -70 mV predominantly probes entry into slow inactivation from closed states, whereas 0 mV drives channel activation and more effectively induces slow inactivation. We therefore chose 0 mV as the primary conditioning potential because it is widely used in conventional slow inactivation protocols and induces slow inactivation more robustly than conditioning voltages at -70 mV. We have added this explanation in Methods section of revised manuscript, please refer to Page 20, Lines 569-571.

      (3) The authors mentioned that the insertion of V5 markedly reduced the PRRT2 protein level; thus, Prrt2-V5 knock-in mice could be considered as PRRT2 knock-down mice. Is there any noticeable difference in phenotype between Prrt2-V5 knock-in mouse and Prrt2-mutant mouse? In other words, is PRRT2 knockdown sufficient to affect neuronal excitability, or is a complete PRRT2 ablation required?

      We thank the reviewer for raising this concern regarding the functional consequences of reduced PRRT2 expression in the Prrt2-V5 knock-in mice. Given that PRRT2 protein levels are markedly reduced in this line, and that cerebellar stimulation-induced dystonia is a characteristic phenotype of PRRT2 deficiency, we tested whether Prrt2-V5 knock-in mice also exhibit this phenotype. We found that electrical stimulation of the cerebellar cortex induced dystonia-like attacks in a subset of Prrt2-V5 knock-in mice. These dystonic behaviors resembled those previously observed in Prrt2-mutant mice, whereas no such behaviors were induced in wild-type mice (Figure 6-figure supplement 1). These findings indicate that a substantial reduction of PRRT2 expression (approximately 80%) is sufficient to impair neuronal function and elicit a disease-relevant phenotype in a subset of animals, supporting the interpretation that the V5 knock-in allele is hypomorphic. We have incorporated these results into the revised manuscript, please refer to Page 9, Lines 265-267; Figure 6-figure supplement 1.

      (4) In Discussion (Page 13, lines 358-361), the authors mentioned a putative interaction between PRRT2 and the Nav channel by modeling, while there is no related data. Please either add modeling data or remove those sentences.

      We thank the reviewer for this suggestion. To avoid over-interpretation, we have removed the statements regarding the AlphaFold-based interaction model from the revised manuscript. We agree that the interaction interface remains to be demonstrated experimentally, and we now discuss this point in the Limitations section. Please refer to Page 16, Lines 465-468.

      (5) Typo: Page 14, line 399, "TMEM232" should be "TMEM233".

      We thank the reviewer for pointing out this typo. We have corrected it in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Mechanistic depth: While the functional data show altered slow-inactivation kinetics, the mechanistic explanation remains superficial. The AlphaFold-based prediction of PRRT2 interaction with DIV-S3 is speculative. The authors should clarify their illustrative rather than evidential intent and avoid over-interpretation.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav interface, including targeted mutagenesis, crosslinking, and structural determination, will be necessary to elucidate the molecular basis of this interaction and its effect on channel gating. Please refer to Page 16, Lines 465-468.

      (2) Separation of trafficking vs. gating effects: Previous studies showed PRRT2 influences Nav trafficking and surface expression. Here, surface expression changes are not systematically quantified. Such an analysis would strengthen the argument that gating effects are not secondary to altered channel abundance or localization.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. We agree that direct analysis of Nav channel surface localization during prolonged depolarization and hyperpolarization would provide stronger evidence to distinguish gating effects from trafficking-dependent mechanisms. However, such experiments are technically challenging in this context: conventional surface biotinylation assays do not provide the temporal resolution required for these rapid protocols, and live-cell imaging approaches to monitor dynamic changes in Nav channel surface expression during slow-inactivation paradigms have not yet been established in our laboratory.

      Although PRRT2 has been reported to regulate Nav channel surface expression in heterologous systems, we consider it unlikely that trafficking is the major determinant of the slow-inactivation effects described here. Slow-inactivation develops on a timescale ranging from tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer timescales (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). We have expanded the Discussion in a revised manuscript. Please refer to Pages 13, Lines 378-388.

      (3) Isoform generalization: Data on other Nav channel subtypes are presented as evidence of a conserved mechanism. However, given tissue-specific expression of PRRT2, these findings may be of limited in vivo relevance. At the very least, additional studies with Nav1.6 should be carried out.

      We thank the review for this suggestion. In response, we conducted new experiments to examine the effect of PRRT2 on Nav1.6 slow inactivation. These results show that PRRT2 promotes entry of Nav1.6 channels into slow-inactivated states and delays their recovery, consistent with its effects on the other Nav isoforms examined in this study. We have incorporated these new data into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      Furthermore, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Pages 14 and 15, Lines 414-416 and 429-430.

      (4) In vivo functional link: The EEG after-discharge threshold assay suggests decreased cortical resilience, but causality between slow-inactivation impairment and hyperexcitability remains indirect. Complementary in vivo recordings would strengthen the physiological link.

      We thank the reviewer for this helpful suggestion. To further link impaired slow-inactivation to the hyperexcitability, we applied a repetitive stimulation protocol in corpus callosum slices, a white-matter region of brain enriched in both PRRT2 and Nav channels. During high-frequency stimulation (e.g., 20 Hz), the amplitude of the compound action potential progressively decreased over the course of the stimulus train. This phenomenon, often referred to as adaptation, reflects activity-dependent reduction in Nav channel availability (Fleidervish et al., 1996; Mickus et al., 1999; Kim et al., 2012). Compared with wild-type mice, Prrt2-mutant mice exhibited less adaptation during high-frequency stimulation, consistent with impaired slow inactivation during repetitive activity, which may contribute to hyperexcitability (Figure 7-figure supplement 2). We have added these results to the revised manuscript. Please refer to Pages 11, Lines 311-322; Figure 7-figure supplement 2.

      (5) Structural interaction: It remains unclear whether PRRT2 binds the α-subunit directly or through accessory proteins. Crosslinking or detergent-solubilization controls of different stringencies could clarify this.

      We thank the reviewer for raising this important issue. We agree that our co-immunoprecipitation data do not distinguish whether PRRT2 associates with the Nav channel α-subunit directly or through other components of the protein complex. To avoid over-interpretation, we have revised the relevant text in the manuscript to remove any implication of direct binding and now describe the result as an association between PRRT2 and Nav channels.

      We have also expanded the Limitations section to note that additional experiments, such as crosslinking and structural studies, will be required to define the interaction interface between PRRT2 and Nav channels. Please refer to Page 16, Lines 465-468.

      (6) Comparisons to other regulators: The paper positions PRRT2 as distinct from FHFs and β-subunits. The data support this, but the discussion could more critically assess whether PRRT2 acts by stabilizing a pore-based inactivated conformation, as suggested for other slow-inactivation modulators.

      We thank the reviewer for this insightful suggestion. At present, relatively few modulators have been characterized in detail with respect to their effects on Nav channel slow-inactivation kinetics. Moreover, even for compounds such as lacosamide, which has been proposed to act as a slow-inactivation modulator, the underlying mechanism remains under debate (Errington et al., 2008; Jo and Bean, 2017). Therefore, in the revised manuscript, we discussed the possible mechanism of PRRT2 in the context of current models of Nav channel slow inactivation.

      Previous studies suggest that entry into the slow-inactivated state involves at least two coupled processes: conformational changes in the voltage-sensing domains and structural rearrangements in the pore region, including the selectivity filter and intracellular activation gate (Catterall et al., 2024; Silva, 2014). During prolonged depolarization, voltage sensors become stabilized in the up-state, while the pore undergoes progressive rearrangements associated with slow inactivation (Balser et al., 1996; Vilin et al., 1999). Thus, mechanisms that further stabilize voltage sensors in the up-state and/or facilitate pore-based inactivated conformations could enhance slow inactivation.

      Within this framework, PRRT2 may enhance slow inactivation by facilitating one or both of these processes, although direct evidence is still lacking. We have incorporated this discussion in relative section of revised manuscript. Please refer to Page 14, Lines 389-404.

      Response references:

      Jo S, Bean BP. Lacosamide Inhibition of Nav1.7 Voltage-Gated Sodium Channels: Slow Binding to Fast-Inactivated States. Mol Pharmacol. 2017 Apr;91(4):277-286.

      Errington AC, Stöhr T, Heers C, Lees G. The investigational anticonvulsant lacosamide selectively enhances slow inactivation of voltage-gated sodium channels. Mol Pharmacol. 2008 Jan;73(1):157-69.

      (7) Behavioral/clinical link: Given the strong human genetics background of PRRT2 disorders, a brief analysis or reference to electrophysiological phenotypes in patient neurons would contextualize the cortical findings.

      We thank the reviewer for this suggestion. Previous studies showed that iPSC-derived excitatory neurons from a patient carrying a homozygous PRRT2 mutation exhibited increased sodium currents and neuronal hyperexcitability (Fruscione et al., 2018). Given that slow inactivation regulates Nav channel availability and thereby influences neuronal excitability, these electrophysiological abnormalities in patient-derived neurons may, at least in part, reflect impaired PRRT2-dependent regulation of Nav channel slow inactivation. We have added this point to the relative section of the revised manuscript. Please refer to Pages 15, Lines 432-437.

      Minor comments

      (1) Figures should include statistical sample sizes (n) and ideally overlay data points rather than only means {plus minus} SEM.

      We thank the reviewer for this suggestion. In the revised manuscript, we present both individual data points and mean ± SEM in the column graphs. For the line graphs, individual data points were not overlaid because of space and readability constraints, and these panels therefore display mean ± SEM only. Sample sizes for each group are provided in the corresponding figure legends.

      (2) The AlphaFold model should be provided as a supplementary figure with confidence scores indicated.

      We thank the reviewer for this suggestion. However, because the predicted Nav1.2-PRRT2 interaction interface has not yet been experimentally validated in our study, we chose to remove the AlphaFold-based model from the revised manuscript to avoid over-interpretation.

      (3) Clarify whether TTX sensitivity was verified in the axonal bleb preparation.

      We thank the reviewer for raising this point. We verified the identity of the sodium currents in the axonal bleb preparation by their sensitivity to TTX, and this information has now been added to Figure 7A in the revised manuscript. Please refer to Page 10, Line 290; Figure 7A.

    1. eLife Assessment

      This important study investigates how surface stickiness shapes whisker mechanics and peripheral neural responses during active touch. The biomechanical evidence that surface stickiness alters whisker mechanics and stick-slip dynamics is compelling, supported by a large and high-quality 3D dataset, while the electrophysiological evidence is solid but limited by a small sample size and insufficient validation of the sticky stimuli. The work will be of broad interest to sensory neuroscientists studying active touch.

    2. Reviewer #1 (Public review):

      Summary:

      This study offers a careful and technically strong look at how surface stickiness changes whisker-surface interactions and how that information reaches peripheral sensory neurons. The authors use 3D whisker tracking to capture bending, twisting, rolling, and tip motion during contact with surfaces that differ in stickiness, coarseness, and position. They show that sticky surfaces, especially silicone, broaden the range of whisker deformation, produce stronger but less frequent stick-slip events, and change firing rates in some trigeminal ganglion neurons. Overall, the study is valuable because it goes beyond standard 2D tracking and shows that out-of-plane motion and roll are important for understanding how whiskers encode texture.

      Strengths:

      The study is technically strong and well motivated. Its main strength is the use of 3D whisker tracking to show that surface stickiness affects whisker deformation in ways that standard 2D tracking would miss, including torsion, roll, out-of-plane motion, and stick-slip dynamics. The authors also connect these mechanical effects to TG activity, providing evidence that stickiness information is available in peripheral sensory responses. Overall, the work expands the study of whisker-based texture sensing beyond coarseness and provides a richer biomechanical framework for understanding tactile encoding.

      Weaknesses:

      The main weakness is that stickiness is not formally defined early in the manuscript, even though it is the central experimental variable. Several methodological choices also need clearer justification or validation, including the use of 2D measures as comparators for torsion and roll, the thresholds used for stick-slip detection, the degree-5 polynomial fit, the reference ROI, and aspects of the 3D surface reconstruction. The neural evidence should also be interpreted cautiously because the TG sample is small, only a subset of units discriminated silicone, and the correlation between strain sensitivity and silicone discrimination is suggestive rather than definitive.

    3. Reviewer #2 (Public review):

      The authors explore the sensation of stickiness from the point of view of whisker exploration and encoding in the trigeminal ganglion. In doing so, they develop methods for 3D whisker tracking to describe stick-specific parameters such as stick-slip rates and strain. Overall, the methods are strong, and the authors present the results appropriately. Overall, I think exploration of the sensation of stickiness is a great question.

      My main criticism is in relation to the chosen stimuli, and I wonder whether the authors may have room to explore more naturally sticky materials and what this may mean for the animal.

      (1) Chosen stimuli for stickiness:

      Four different materials are used, with the aim of presenting animals with graded measures of stickiness. The results show that silicone stands out against the others; it's less clear whether the intermediate textures (Delrin and resin) may be truly intermediate in stickiness.

      I wonder if the stimuli chosen were truly representative of the aim of providing a gradient of stickiness. Did the materials differ in other features, such as surface temperature, texture, etc., which could explain some results? The authors discuss this in terms of coefficients of friction and how these estimates are not quantified in relation to whiskers themselves.

      Measures of stick-slip and strain with silicone vs other materials make intuitive sense. Could the authors add additional naturally sticky stimuli to exemplify the results? For example, adhesive, glue, or a sugary substance.

      (2) Tracking methods and quantification:

      The 3D tracking methods, which incorporate whisker twists, strain, and other fine features of whisker exploration, present an advance in terms of analysis of how whiskers may explore more complex, natural features of environments. The analyses and quantifications are all solid and robust. The technical approaches are well-prepared to take the work a step further in terms of stimulus choice.

      (3) Peripheral coding of stickiness:

      The authors report that some units respond preferentially to whisking on silicone and that this has to do with strain on the whisker. Is there a possibility to understand the nature or anatomy of these units and why they might be preferential for the sticky sensation? Can the location in the follicle be assigned? And/or would the methodology allow for assignment of where the specifically sticky-tuned units project centrally?

      (4) Relationship to natural stimuli:

      A piece missing from the paper is more discussion and exploration of why stickiness may be important for sensory coding, as well as potentially more naturally sticky stimuli. One could imagine that a mouse navigating the world could find stickiness attractive, if it were a source of sweet food, for example, or it could potentially be a sensation the animal prefers to avoid. Stickiness could also indicate contamination or a sticky trap, to be avoided. If the authors are able to add naturally sticky stimuli, the whisker exploration and encoding could potentially provide further cues towards the valence of stickiness for mice.

    4. Reviewer #3 (Public review):

      This paper tackles an underexplored dimension of whisker-based texture sensing: while surface coarseness encoding has been extensively characterized in rodents, the mechanical and neural basis for stickiness sensing has not previously been examined. The authors make two intertwined contributions that together represent a substantial advance: a methodological one - a 3D whisker tracking pipeline operating at 4000 fps, capable of capturing torsion, roll, and out-of-plane whisker motion - and a scientific one - a first characterization of how whisker mechanics and primary trigeminal afferent responses differ between surfaces of high and low stickiness. The work is technically solid, the dataset is large, and the question is well motivated both by the multidimensional nature of tactile texture perception and by the practical advantages of the whisker system for studying touch mechanics.

      Strengths.:

      The 3D tracking system is a timely advance over existing tools, particularly in its handling of non-planar whisker shapes and the full automation required for the sub-millisecond resolution needed to detect stick-slip events. The mechanical dataset is extensive. The finding that whisking against silicone expands the sampled whisker strain space and produces stronger but less frequent stick-slip events is clearly demonstrated and internally consistent with the proposed mechanism of greater strain accumulation before frictional release - a physically intuitive result. The open release of the tracking code considerably increases the value of this work to the broader community.

      Weaknesses:

      A few aspects of the paper, if sharpened, would considerably strengthen the evidence and the clarity of the conclusions.

      The central claim - that "stickiness information is available to the whisker system" - does not capture the precision of what the paper demonstrates. As stated, the finding is close to guaranteed: any variation in surface friction will produce some change in whisker mechanics, so the presence of mechanical differences between materials is expected rather than surprising. The more valuable question the paper is well positioned to answer is which specific dimensions of the whisker mechanical response are most informative about surface stickiness. The paper reports effects on strain distribution breadth, stick-slip amplitude, and stick-slip rate, but does not synthesize which of these - or which sub-dimensions (bending, twisting, or rolling) - carry the most discriminating information. Identifying the salient dimensions of the mechanical response and relating them to the proposed frictional mechanism would sharpen the paper's conclusions substantially.

      A related but distinct limitation is the absence of direct force measurements during whisker-surface contact. The authors acknowledge this openly, and I recognize it is not easily remedied within the current experimental setup. It does, however, constrain interpretation: without knowing the actual forces generated at the whisker-surface interface, the assumed stickiness ordering of the tested materials cannot be validated, and - importantly - the relative contribution of surface friction and material compliance to the observed mechanical differences cannot be determined. This is an important direction for future work in this area.

      The paper argues carefully that 2D tracking is insufficient for capturing the full mechanical picture of whisker-surface interactions, and the figure currently in the supplementary material (Figure S2) makes this case convincingly through multiple analyses. This argument is the core justification for the paper's methodological contribution and deserves a place in the main manuscript. Furthermore, while the mechanical case for 3D over 2D tracking is well made, it has not yet been tested at the neural level: the regression model used to predict neural firing incorporates 3D variables, but its performance is not compared against an equivalent model restricted to 2D variables. Such a comparison would directly demonstrate whether torsion and roll - the signals inaccessible to 2D tracking - carry neural predictive value, and would elegantly unite the paper's methodological and scientific contributions.

      Finally, the three-dimensional plots in Figure 3 are the paper's primary representation of its main mechanical result, and there is a real opportunity to make them considerably more informative. The whisker deformation probability distributions (panel B) are rendered in 3D from a single viewing angle, making it difficult to assess the shape or anisotropy of the distributions - and in particular to see which dimensions expand most for silicone relative to the other materials. This is precisely the information needed to identify the most salient dimensions of the stickiness signal, and two-dimensional representations would make it directly readable.

    1. eLife Assessment

      This study presents a useful compendium of triangulated single-cell eQTLs, Mendelian randomisation and colocalization of genetic signals in prostate cancer datasets. Biological interpretation in the context of the aging prostate gland, the tumour microenvironment and immune cell specificity is incomplete, so this study is a starting point for further study, and would require validation of the resulting putative causal genes.

    2. Reviewer #1 (Public review):

      Summary:

      Using Mendelian randomisation on available GWAS data, the investigators identified eGenes associated with prostate cancer and applied the data to define relevant immune cell types involved. Additional analysis was performed to explore potential candidate targets and agents from licensed medicines.

      This is an interesting approach as the investigators have expertise in other research fields, applied here to prostate cancers. The use of three different datasets is significant, and the approach to further analyse implicated eGenes in drug target analysis is relevant and timely.

      A particular strength is taking putative genes from Mendelian randomisation analysis to target and potential drug agents.

      Some aspects of the study would need to be clarified to enable interpretation of the findings in the context of the prostate gland and prostate cancers: expanding the descriptions of the supporting Supplementary Data and Tables, explanations of the analysis for the general reader, and clarification of the selection of eGenes (Figure 5).

    3. Reviewer #2 (Public review):

      Summary:

      This study integrates bulk and single-cell transcriptomic-derived eQTLs from two separate consortia (PRACTICAL and Finngen) to identify immune-cell-specific therapeutic targets in prostate cancer. Mendelian randomization and Bayesian colocalization have been used to produce druggable eGene modules through STRING and DrugBank.

      This is an interesting study that is attempting to address risk-associated, immune-specific transcriptomic repertoires in prostate cancer. It is knitting together concepts of drug repurposing and prostate cancer immunogenicity. This is an entirely computational study, which would benefit from some wet lab experimental validation.

      It is very tricky to attribute cell-type-specific responses, especially when the majority of genes involved represent cytoskeletal or stress responses, which are ubiquitous throughout the prostate microenvironment. This point is relevant for the drug repurposing section: if these drugs are targeting immune cell-specific repertoires, what would the response be of the entire environment? It would be useful to contextualize the validity of each proposed therapy in a specific prostate cancer context and the involvement of AR antagonism or radiotherapy.

      Strengths and limitations of this study:

      Strengths:

      This is a scientifically interesting and potentially impactful study, particularly in its attempt to integrate immune-cell-specific transcriptomics, causal inference, and drug repurposing in prostate cancer. The methodology is well described, and the data (albeit limited) are well analyzed.

      Limitations:

      The central weakness is the overstatement of the conclusions regarding immune-cell-specific causality, without sufficiently contextualizing the biological meaning of the findings.

      Highlighted genes, such as LMNA, XBP1, histone-related genes, and stress-response markers, are ubiquitous regulators involved in fundamental cellular processes, including ageing, unfolded protein response (UPR), integrated stress response (ISR), chromatin remodeling, proliferation, and metabolism. It is unclear whether these signatures truly represent immune mechanisms, or instead reflect broader inflammatory and age-associated biology expected within an ageing glandular organ such as the prostate.

      Immune cell identity alone may not be sufficient to infer biological relevance because immune state characterization (e.g., exhausted versus functional T cells, or distinct macrophage/myeloid phenotypes) is largely absent from the current analysis. The assertion that specific immune populations are correlated with prostate cancer susceptibility is probably an overstatement unless the nature of these cells can also be characterized.

      The interpretation of "causal variants" is not always specified, i.e., what phenotype is being associated: prostate cancer susceptibility, recurrence, progression, or treatment response (e.g. is there direct causality from immune-cell variants to prostate cancer?).

      Overall, there is a need for stronger biological and translational contextualization: how do the identified pathways relate to ageing-associated inflammation, PIN, microbiome-driven inflammatory changes, and stress-response biology in the prostate gland? While the manuscript identifies network hubs and enriched pathways, it often stops short of explaining what these modules biologically represent or how they may influence prostate cancer development, progression, treatment resistance, or immune evasion.

      There are additional publicly available spatial transcriptomic or single-cell datasets which could be used to validate whether the purported immune-cell-specific genes are genuinely enriched in immune populations adjacent to tumour cells. In the drug repurposing analyses, the current study does not explicitly handle prostate cancer subtypes such as HSPC, CRPC, NEPC, or DNPC and co-treatment with androgen receptor antagonism or radiotherapy.

    1. eLife Assessment

      This useful study explores how macrophage cell-cycle state may influence endocytosis, Mycobacterium tuberculosis uptake, and the intracellular stress experienced by bacteria. While the question is interesting and the experimental approach has promise, the evidence for the central claim that endocytic capacity is specifically regulated by cell-cycle stage is incomplete. The main concern is that fluorescence-based sorting and total-fluorescence measurements likely covary with cell size, so the reported phenotypes could reflect biomass accumulation or other cell-cycle-associated changes rather than endocytic capacity as the causal determinant. As a result, whilst the study raises a hypothesis that is of importance, additional controls are required before the proposed mechanism can be considered well supported.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript- "Cell cycle-dependent variation in endocytosis drives phenotypic diversity in M. tuberculosis" by Subhash et al. demonstrates how host cell heterogeneity shapes intracellular pathogen phenotypes. The central and novel finding of this study (G2-phase cells have higher endocytic capacity and harbour more oxidised Mtb) highlights that a host cell cycle (interphase-driven) changes in endocytic capacity regulate bacterial redox states.

      Strengths:

      Overall, the study is well-executed and conceptually rich, establishing a causal link between host cell cycle progression, endocytic heterogeneity, and M. tuberculosis phenotypic diversity.

      The combination of multiple modalities, including live-cell imaging, flow cytometry, scRNA-seq, and redox-sensitive bacterial reporters, supports these findings and substantially strengthens the biological relevance of the work.

      The writing is generally clear, and the figures are well-organised.

      This work will be of interest to readers across cell biology, microbiology, and infection biology

      Weaknesses:

      However, several central claims are only partially supported, the mechanistic depth is limited, and several experimental and analytical concerns need to be addressed.

      Major Comments:

      (1) The authors demonstrate a correlation between the G2 phase and elevated endocytic capacity. However, the mechanistic link (upstream molecular mechanism) between the cell cycle and endocytic upregulation remains largely unaddressed. The authors speculate that membrane biogenesis during volumetric expansion may drive increased endocytosis and note that lipid biosynthesis genes are upregulated in high-endocytic cells. It would substantially strengthen the paper to test this directly, by examining whether inhibition of lipid biosynthesis (e.g., with fatostatin or cerulenin) selectively reduces the G2-associated increase in endocytic capacity. Alternatively, cyclin-CDK axis perturbations (e.g., CDK1 inhibition with RO-3306 to specifically block G2/M entry) could be used to ask whether cells arrested in G2 maintain elevated endocytosis, helping distinguish cell-cycle-position-dependent from cell-cycle-progression-dependent effects.

      (2) The current data show a clear association between high endocytic capacity and more oxidised Mtb, and the authors (consistent with their prior work) hint at lysosomal delivery as the likely mechanism. However, direct evidence for this in the current paper is limited. An experiment examining phagosomal pH or lysosomal fusion (e.g., using a pH-sensitive reporter or lysotracker) specifically in high- and low-endocytic-capacity cells after infection would help confirm this.

      (3) Temporal resolution of Mtb redox dynamics. The plasticity experiment (Figure 6C-D) is elegant and shows that Mtb redox states revert as host cells divide and daughters enter G1. However, the experiment compares day 0 and day 3 post-sorting, which spans multiple cell divisions. While a finer time resolution (spanning 24h) would establish the causal relationship, the authors could discuss the possibility and consequences of multiple cell divisions between day 0 and day 3 used in the present study.

      (4) Relevance of G2 percentages in differentiated macrophages. In Figure 7 and Supplementary Figure S7, only 4.4-5.7% of THP-1-derived macrophages and 5.7% of BMDMs are in G2. While the authors demonstrate statistically significant differences in Mtb redox states between G1 and G2 macrophages, the biological significance of such a small G2 fraction in a non-dividing population deserves discussion specifically with respect to: a) Are these cells re-entering the cycle? b) Is the G2 designation capturing a distinct functional state rather than active cycling? The authors should include additional markers (e.g., phospho-histone H3 for mitotic cells or BrdU incorporation to test for active S-phase) to characterise this population and clarify its identity and origin in differentiated macrophages, thereby meaningfully informing interpretation.

      In conclusion, this is an important mechanism-driven study that highlights an important link in host-driven bacterial phenotypic heterogeneity. The experiments are thorough, the model is well-supported, and the study has implications for infection biology.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors utilize a combination of techniques to show that macrophage endocytic capacity is partially dictated by cell-cycle stage, that Mycobacterium tuberculosis (Mtb) more readily infects macrophages that are in G2/M -phases, and that bacteria that are internalized by macrophages at different stages of the cell-cycle experience different levels of intracellular stress (as reported by the redox state of the bacteria). Furthermore, the authors provide evidence that terminally differentiated macrophages retain memory of the cell-cycle stage that they were in prior to differentiation, at least in the context of endocytic capacity.

      This work provides evidence for the growing idea that fundamental heterogeneity in both host and bacterial organisms can alter the host-pathogen relationship in important ways. However, based on the current data, I am not convinced that the manuscript establishes endocytic capacity as the causal link between macrophage cell-cycle stage and bacterial state. The main issue is that fluorescence-based sorting for cell-cycle stage is likely to covary with cell size. Larger cells, including those later in the cell cycle, may be more likely to fall into the "high" fluorescence gate, while smaller cells may be enriched in the "low" population. Therefore, the observed phenotypes may still be cell-cycle-associated, but the causal determinant could be a correlated feature of cell-cycle progression rather than endocytic capacity itself. This is a significant caveat because nearly all the data, including the live-cell imaging following individual cells, rely on 'total' fluorescence, which will scale strongly with cell size.

      If the authors' conclusion that endocytic capacity is cell-cycle regulated holds true after appropriate controls, this would significantly advance our understanding of the causal interplay between host cell-cycle state, endocytosis, and Mtb physiology. However, an alternative interpretation is that the observed differences in Mtb uptake and bacterial redox state are associated with cell-cycle stage but are not caused directly by differences in endocytic capacity. For example, they could instead reflect other cell-cycle-linked changes in macrophage physiology, such as cell size, intracellular volume, metabolic state, or some other mechanism important for Mtb pathogenesis. If the authors find that their data are best explained by cell-cycle stage independent of endocytic capacity, this would still represent an important advance. However, in that case, the manuscript should clearly distinguish the association with cell-cycle state from the downstream effector mechanisms, which would remain to be determined.

      Strengths:

      The authors utilize various macrophage models for their studies, which is important considering the variability in macrophage behavior, as well as the growing evidence that differences between mouse and human macrophages are relevant for Mtb infection.

      Weaknesses:

      The most important caveat is the covariance between fluorescence-based reporters and cell size. This concern applies to both the sorting experiments, which directly measure total fluorescence, and the time-lapse microscopy experiments, in which the authors show total fluorescence rather than mean, area-normalized fluorescence in Figure 3C. This could be explained by biomass accumulation alone, rather than by a specific cell-cycle-dependent increase in endocytic capacity. Without distinguishing total signal from concentration or activity per unit cell area/volume, it is difficult to conclude that endocytosis itself is regulated by cell-cycle stage rather than simply scaling with cell size.

      Although the authors provide some evidence that the mean GFP intensity, which more closely reflects concentration, differs between the sorted populations in Figure 3B, they do not report statistics for this comparison. Moreover, this control is not carried through the rest of the manuscript, including in key experiments such as Figure 2B. As a result, it remains difficult to determine whether the observed differences between "high" and "low" populations reflect cell-cycle state specifically or instead reflect differences in total reporter fluorescence driven entirely by cell size.

      The evidence for cell-cycle-dependent effects would be more convincing if the authors included additional controls. For example, they could:

      (1) Plot both mean GFP intensity and total GFP intensity in Figure 3B, ideally alongside an unrelated fluorescent reporter that does not vary across the cell cycle. This would help distinguish changes in reporter concentration from changes driven by cell size or total fluorescence.

      (2) Sort cells based on an unrelated fluorescent marker and test whether the same phenotypes - infectivity, dextran uptake, bacterial redox state, etc. - differ between high- and low-fluorescence populations. If these phenotypes are specific to the cell-cycle reporter and not observed with an unrelated marker, this would strengthen the conclusion that the effects are linked to cell-cycle state rather than to fluorescence intensity, cell size, or sorting artifacts.

    1. eLife Assessment

      This important submission from Ambler and colleagues brings new insights into how torpor conditions may confer resilience in cases of cardioprotection. It has novelty, which can be enhanced by additional in vivo support. The study is backed by solid evidence, and represents a unique interoceptive mechanism of interest.

    2. Reviewer #1 (Public review):

      Summary:

      Torpor can be induced by chemogenetic activation of the medial preoptic area. This activation leads to protection from myocardial infarction in an isolated heart preparation despite normalization of the ambient temperature, thus, in principle, uncoupling hypothermia from torpor-induced neuroprotection. Putative pathways of protection are suggested by proteomic studies.

      Strengths:

      (1) Elegant strategy for inducing torpor in rats.

      (2) Appropriate controls for verifying the neuron transducer.

      (3) Cardiac protection is significant and appears independent of hypothermia.

      (4) Interesting omic strategy to begin to find established and novel pathways mediating organ autonomous torpor-induced protection.

      Weaknesses:

      (1) The study would benefit from using inhibitory chemogenetics of the same neurons to demonstrate that this might make cardiac response to ischemia worse.

      (2) Infecting an area of the brain not known to be involved in torpor would be a useful control.

      (3) In vivo cardio protection seems essential as the validation of the strategy requires support that is in the intact animal.

      (4) The assumption that the positive effects of torpor are mediated via a phosphoproteomic change rather than a translational or transcriptional control mechanism is not established.

      (5) A 40 percent reduction in infarct size may work for genetically identical rats with no co-morbidities, but is unlikely to be significant enough to weather the variability that emerges in humans because of these differences and more. The question is not what the mechanism is, but how do we make it more robust? Overall, this is at best a preliminary data set that requires more experiments to deliver on its immense promise.

    3. Reviewer #2 (Public review):

      Summary:

      Elley and colleagues induced a synthetic torpor-like state in rats (a non-hibernating species) by chemogenetically activating neurons in the medial preoptic area of the hypothalamus. They show that this state substantially reduced cardiac infarct size in an ex vivo ischemia-reperfusion model. They further report that protection persisted when ambient temperature was raised to prevent hypothermia, and used exploratory phosphoproteomics to identify candidate cardioprotective signaling pathways.

      Strengths:

      This is the first demonstration that a torpor-like state is cardioprotective in a species that does not naturally enter torpor, which meaningfully advances the potential clinical utility of synthetic torpor. The experimental design is logical, and the controls are generally appropriate. The characterisation of the responsible neuronal population using ISH against QPLOT markers adds mechanistic depth and supports the cross-species conservation argument. The phosphoproteomic analysis, though exploratory, generates plausible and biologically coherent hypotheses grounded in the hibernation literature.

      Weaknesses:

      The primary weakness is that the central conclusion - that hypothermia is not necessary for cardioprotection - exceeds the evidence. The thermoneutral groups were not demonstrably normothermic (36.4 vs 37.05{degree sign}C, p=0.44 with n=6), core temperature telemetry was absent in the majority of control animals contributing to the infarct endpoint, and the decisive test, i.e., a correlation between individual nadir temperature and infarct size, was never performed. Additional weaknesses include the absence of sex-stratified analysis despite known estrogenic contributions to torpor

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Elley and colleagues describes experiments on the effects of synthetic torpor on ex vivo heart ischemia. The key aspect of the study was the use of viral-vector mediated manipulation of the hypothalamic medial preoptic area (MPA) in rats. They used AAV-CaMKIIa-hM3D(Gq). The authors report that chemogenetic activation of the MPA prior to an ex vivo heart ischemia-reperfusion insult induces cardio protection against infarct size that is independent of prior in vivo hypothermia. Phosphoproteomic analysis of cardiac tissue suggested changes in cell survival and death pathways.

      Strengths:

      This study has important strengths. The idea is novel. The experimental design is appropriately rationalized and fascinating. The manuscript is written and presented concisely.

      Weaknesses:

      The study has important weaknesses in the experimental design and validation of the model.

      (1) The study is based on the use of a DREADD-designed viral vector (AAV-CaMKIIa-hM3D(Gq) -mCherry) that is activated by 2 mg/kg IP injection of CNO. The rationale is to putatively activate the MPA. The authors show no evidence for chemogenetic activation of neurons in the MPA. This could be done using a variety of different approaches, even phosphoproteomics.

      (2) The stereotaxic injections are difficult to precisely and locally place, particularly bilaterally. Figure 2F is only a schematic. It would be better to show actual low magnification brain sections (bregma +0.12 to -0.48) from a representative rat to show the placement of the AAV.

      (3) The control rats were injected with AAV-CaMKIIa-EGFP. Why was EGFP used instead of mCherry for the control?

      (4) Ideally, a mutant non-activatable variant of AAV-CaMKIIa-hM3D(Gq) should have been used for a better control.

      (5) The authors should comment on whether there is any neurotoxicity in the MPA associated with the forced AAV expression of hM3D-Gq.

      (6) Is there any inflammatory pathology seen in the MPA with AAV transduction?

      (7) There are no experiments to show that the systemic torpor is specifically associated with the MPA region. Experiments should be done with injections of AAV-CaMKIIa-hM3D(Gq)-mCherry placed in other brain regions, for example, the nearby nucleus accumbens.

      (8) The mapping of the distribution of neurons responsible for synthetic torpor is not mechanistic enough and is not directly to the point. While excitatory and inhibitory markers are examined, a more interesting and deeper approach would have been to use glutamate receptor antagonists to manipulate the torpor response.

      (9) The ischemia and reperfusion aspects of the Lagendorff method need to be clarified. The isolated hearts are already ischemic after their removal from the rat. The reperfusion aspect is caused by reflow of blood to generate oxidative stress, but in the ex vivo model, is there really reperfusion injury?

      (10) The authors show that whole animal oxygen consumption is reduced in the torpor state. The measurement is crude and most likely reflects the inactivity of the animal's skeletal muscle in the torpor state. A more relevant and direct experiment would be to do oxygen consumption (or Seahorse) assays on extracts of the isolated hearts.

      (11) The authors report that the synthetic torpor induces bradycardia. There is no follow-up on this important observation. The MPA-heart connection is not analyzed. (A) Is the link through cardiovascular centers in the brainstem? (B) Is the torpor-induced bradycardia mediated through increased parasympathetic or decreased sympathetic autonomic tone? Pharmacological experiments could also be done.

    1. eLife Assessment

      This study addresses a recent discovery by others that electroconvulsive therapy (ECT) generates seizure activity and spreading depolarization (SD), reflected in large increases in calcium, which can be followed by imaging calcium fluctuations in neurons. This work is useful. However, the evidence to show that SD, rather than seizures, confers the neuroplastic and other therapeutic effects of ECT is incomplete.

    2. Reviewer #1 (Public review):

      The work corroborates the idea, recently suggested by Rosenthal et al. (2025), that spreading depolarization is involved in the mechanisms of electroconvulsive therapy. Using a mouse model of electroconvulsive therapy and various sophisticated approaches to visualize cortical activity, the authors provide an extensive description of traveling calcium waves induced by electroconvulsive stimulation. The study confirms that the calcium events have properties typical of cortical spreading depolarization and seeks to show that the calcium/SD waves mediate therapeutic and neuroplastic effects of electroconvulsive therapy. The authors find that after electroconvulsive stimulation associated with calcium/SD waves, Fos expression increases widely; in the cortex, this increase is localized to the hemisphere affected by calcium waves. They show that some EEG predictors of the beneficial effects of electroconvulsive therapy correlate with the occurrence of calcium/SD waves. Despite the solid methodology and the study's interesting, its conclusions are not fully supported by the data.

      In particular:

      (1)The title of the paper claims that "electroconvulsive stimulation drives cortical spreading depolarization dependent immediate early gene expression". However, immunohistochemical staining shows that Fos expression increases not only in the cortex but also in many subcortical regions, including the hippocampus and amygdala (Figure 5A). Really, conventional electroconvulsive therapy stimulates nearly the entire brain volume and induces generalized seizure activity that can trigger SD not only in the cortex but also in other brain sites. Therefore, regions beyond the cortex can also drive the effects of electroconvulsive therapy. Next, the authors use Fos staining as a marker of neuronal plasticity. However, Fos is also a marker of preceding neuronal activation. As electroconvulsive stimulation, seizures, and SD are associated with high neural activity, it is unclear whether the observed Fos upregulation results from the prior activation or heralds the subsequent plastic changes. Other markers of neuroplasticity (e.g., BDNF) should also be examined.

      (2) Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy. Cortical SD is also tightly coupled with suppression of neuronal activity in affected regions. Although the authors report that postictal suppression is stronger after stimulations with cortical SDs than without SDs, the cortices affected (ipsi) and unaffected (contra) by unilateral cortical calcium/SD events exhibit identical suppression (Figure 6F). The result contradicts established knowledge in the field. If the calcium events are cortical SDs, they should induce EEG suppression only in the affected hemisphere.

      (3) The study states a beneficial role of calcium/SD waves in ECS effects. However, SD alters numerous aspects of brain function, leading to a range of effects that can underlie side effects as well. Assessment of the behavioral effects of stimulation with and without calcium/SD waves can help clarify the issue.

      The results of the work suggest that cortical SD can contribute to electroconvulsive therapy-related mechanisms and help to optimize the stimulation parameters to achieve maximal therapeutic effect.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses the question of mechanisms underlying the therapeutic effects of electroconvulsive therapy (ECT). Clinical efficacy of ECT in major depression (and other disorders) is well established and has often been assumed to be a direct consequence of seizure activity generated by the current application. However, as the authors point out, this explanation is unsatisfactory. A recent study (Rosenthal et al., 2025) provided evidence that ECT generates a wave of cortical spreading depolarization (CSD) in mice, and initial evidence that similar events were generated in patients undergoing ECT. Based on their observations, Rosenthal et al. proposed that CSD, rather than seizure, may engage plasticity mechanisms that contribute to the brain's clinical response to ECT. The current study adds to that prior work by reporting other consequences of CSD, in addition to sustained Ca2+ elevations. The current study also links EEG characteristics immediately following the ECT with the likelihood of generating a CSD, which can help optimize ECT parameters.

      Strengths:

      An important research topic, linking a large set of rodent studies with a limited clinical EEG data set.

      The data acquisition and analyses appear to be of very high quality, and the main results are well illustrated.

      Association between EEG characteristics linked to good clinical outcome matched by mouse EEG data linked to CSD.

      Characterization of multiple consequences of CSD following ECT in the mouse brain.

      Weaknesses:

      The main characterization of CSD propagation comes from GCaMP Ca2+ measurements, as previously reported (Rosenthal et al., 2025). That prior study also provided key electrophysiological evidence of CSD with a DC shift after ECT in mice (supplemental data). Given the prior evidence for ECT-CSD, the additional measures shown in the current manuscript are fully expected. Thus, the 2-photon imaging of Ca2+ elevations following CSD (Figure 4) is consistent with prior 2-photon imaging studies of CSD, and the complex hemodynamic and pH changes are expected to contribute to propagation of EGFP fluorescence changes (Supplemental Figure 5). These data are well presented, but, contrary to the results section here, these results appear confirmatory rather than necessary to build a case that the key event generated by ECT is a CSD.

      The authors state that "our conclusion that CSD is the primary driver of plasticity is based on its role in driving Fos expression" (line 472). Related to the point above, there is already a very well-established literature showing that CSD leads to rapid and robust Fos expression in rodent cortex, so this is fully consistent with prior work. The prior work, CSD-fos work, should be summarized and/or cited more clearly in the manuscript. Showing that Fos increases only in the hemisphere where there is a large CSD-Ca2+ wave is a clear demonstration of this. While Fos increases can certainly be well linked to plasticity in some experimental paradigms, the implication that Fos increases underlie CSD-induced plasticity and possibly therapeutic effects of ECT is not appropriate. Fos increases after CSD are a reliable marker of the very strong neuronal activation that occurs, but Fos increases are not specific for plasticity and can be activated by challenges that do not generate synaptic plasticity. A range of other gene expression changes have been identified with CSD and may contribute to adaptive plasticity; these could be mentioned alongside speculation about Fos. To support the main conclusions of this paper about CSD driving plasticity via Fos, Fos knockout or knockdown studies are needed, as has been used in prior plasticity studies.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript combines widefield calcium imaging, electroencephalography, 2-photon imaging, and immunohistochemistry in mice to re-demonstrate that electroconvulsive stimulation (ECS) induces a seizure followed by cortical spreading depolarization, as previously shown. The putative novel finding - which is not unexpected - is that ECS is also correlated with increased expression of the immediate early gene cFOS, although this has also been shown previously. The authors speculate that CSD drives cFOS expression, which might contribute to the therapeutic effects of ECT; however, experiments performed do not provide causal evidence for this hypothesis. Instead, the authors use expression of cFOS - a nonspecific activity-dependent gene induced in various pathological and non-therapeutic contexts - as a proxy for plasticity and/or therapeutic effect. Hence, overall, the significance of the findings is limited and primarily serves to replicate prior work, with the evidence evaluated as incomplete.

      Strengths:

      The experiments are generally well executed from a technical perspective.

      Main Weaknesses to be addressed in revision:

      (1) The main findings of this paper are replication experiments of prior work, and thus, the novelty and significance of this manuscript are relatively limited.

      - It is already known that the mean frequency of ECT-induced seizures decays between peak and offset in humans (Stuiver et al. Clin Neurophysiol. 2026 Jan:181:2111439. doi: 10.1016/j.clinph.2025.2111439) and mice (Murakami et al. J Pharmacol Sci 2008 Jan;106(1):78-83. 10.1254/jphs.FP0071453), which the authors re-demonstrate in Figure 1.

      - It has already been demonstrated that ECT in mouse models induces lateralized CSD waves in a manner that depends on stimulation parameters and the initial evoked response during stimulation (Rosenthal et al. Nat Comm. 2025 May 18;16(1):4619. doi: 10.1038/s41467-025-59900-1); the authors replicate this in Figures 1, 2, 3, 6.

      - It is already widely established that EEG and calcium signals are highly concordant in mouse brain physiology, as shown in Figure 1. It is already known that CSD propagates from supragranular to granular and infragranular layers (Zakharov et al. Epilepsia. 2019 Dec;60(12):2386-2397. doi: 10.1111/epi.16390) as shown in Figure 4.

      - It is already known that CSD waves induce cFOS expression (e.g., Dell'Orco et al. Front Cell Neurosci. 2023 Dec 14:17:1292661. doi: 10.3389/fncel.2023.1292661; Hermann and Hossman. Neuroscience. 1999 Jan;88(2):599-608. doi: 10.1016/s0306-4522(98)00249-8) as the authors replicate in Figure 5.

      Minimally, the authors should revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field. There is limited innovation in re-demonstrating that these events are seizures and that they involve spreading depolarization.

      (2) The authors frame their hypothesis that CSD could be a potential mediator of the therapeutic effects of ECT, but they do not measure therapeutic effects or directly test this hypothesis. The principal advancement of the paper is showing that ECT-induced CSD triggers hemisphere-specific cFOS expression as a proxy of plasticity. However, it is already known that CSD induces cFOS expression (as noted above). The observation that cFOS expression was induced only by CSD, not by the initial seizure, is likely a byproduct of the greater activity induced by CSD than by seizure. cFOS expression is nonspecific to plasticity or therapeutic effects and can be triggered by many non-therapeutic interventions. The cFOS data thus do not meaningfully measure therapeutic plasticity. The authors also selectively cite references suggesting that EEG metrics such as seizure duration predict positive therapeutic outcomes, but this link is controversial and not well established in the clinical literature.

      Minor Weaknesses:

      (3) For the n=3 mice used for concurrent 2P imaging with microprism implant, these animals also had ChrimsonR co-expression, but there are no optogenetic studies described in this paper, which is confusing. Yet, this co-expression introduces a significant confound, as GCaMP6 emission (525/50nm band in this study) will overlap substantially with the ChrimsonR excitation spectrum. Thus, the fluorescence emission used to image these neurons may be optogenetically activating them at the same time. Please explain.

      (4) Incision of the cortex for implantation of a prism is a significant cortical injury that likely induces CSD instantaneously and may change the propensity for CSD in subsequent recordings. Please comment on this limitation and address how much time elapsed after surgery before imaging.

      (5) Method details are missing or insufficiently described for location, titer, and injection strategy for 2-photon experiments.

      (6) Given the wide range of parameters used for ECS in mice and ECT in humans, the authors should provide tables for what stimulation parameters were used for each recording. These protocols were chosen manually rather than randomly or systematically, which introduces confounding factors into analyses that use parameters as an independent variable.

      (7) While much of the cFOS staining after unilateral CSD shows hemisphere-specific asymmetry, several regions (piriform cortex, amygdala, thalamus) do appear to have bilateral cFOS expression. Please comment on this.

      (8) The discussion states: "If CSD accounts for plasticity effects, triggering a CSD in a non-seizure context may be sufficient to elicit therapeutic effects. This is supported by the clinical success of ultra-brief stimulation treatments that do not cause seizures, such as rTMS with accelerated protocols, which achieves treatment efficacy on par with ECT for major depressive disorder". Are the authors implying that TMS induces CSD? What evidence supports this idea?

      (9) This statement - "Assuming psychosis is the result of thalamocortical coupling that is too weak in frontal areas of the cortex" (lines 583-585) - may be overly speculative.

    1. eLife Assessment

      This important study establishes a robust live-imaging toolkit to characterize excitatory and inhibitory synaptic dynamics during neuronal development, advancing our mechanistic understanding of synaptic homeostasis and neural circuit maturation. The core findings clarify how stable E/I balance is maintained despite persistent synaptic turnover, with broad implications for developmental neurobiology and neurodevelopmental disorders. The methodology and quantitative data are convincing and well validated, and this work represents a significant advance that will be of significant interest to researchers in synaptic biology, cell imaging and neuroscience.

    2. Reviewer #1 (Public review):

      [Editors' note: all three reviewers confirm that all initial concerns have been fully resolved through comprehensive revisions and supplementary analyses.]

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      Comments on revised version:

      The authors have addressed all my questions/comments. No further questions for this manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomato-Gephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HA-Homer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Comments on revised version:

      The authors addressed all my questions and comments. Their edits have made this paper significantly stronger. I believe that this is an important paper for the field.

    4. Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      Comments on revised version:

      The authors have fully addressed my concerns, and this is now a strong manuscript for the synaptic field.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this valuable study, the authors developed long-term imaging tools to simultaneously monitor the temporal and spatial dynamics of excitatory and inhibitory synapses and reported that excitatory and inhibitory synapses need to develop synergistically during synaptogenesis to maintain balance. While the analysis and quantification of the imaging data are incomplete, there is convincing evidence that the developed tools are feasible. If these tools can function stably in vivo, their applications will be much broader.

      We have completely overhauled our analysis and quantification methods and generated custom-made drift correction and tracking pipelines. Also, we have tested these tools ex vivo.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      We thank the Reviewer for their positive feedback and careful review of our study.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      We also thank the Reviewer for their insights and suggestions. We have added discussion of this important point to the Discussion section. Furthermore, we have tested the applicability of our tools ex vivo (new Figures 1, 4, and 6). While using these tools in vivo for live imaging is the eventual goal, we started in a reduced culture system given the relative simplicity. Our current study now provides a framework for future experiments applying these approaches in more complex in vivo systems.

      Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomatoGephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HAHomer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      We thank the Reviewer for their positive assessment of our study.

      Weaknesses:

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Minor weaknesses of the manuscript include:

      (1) The lack of a characterization of endogenous Homer1-positive excitatory synapses using TKIT.

      We attempted to perform live imaging of endogenous Homer1-positive synapses using the TKIT approach by tagging endogenous Homer1 with mClover3 but encountered low signal/noise while live imaging. This prompted us to focus our current study on live imaging endogenous Gephyrin. Future studies using more robust tags (e.g. StayGold, HaloTag) for TKIT tagging of endogenous Homer1 will likely help circumvent this issue.

      (2) Discussion about other approaches to study excitatory and inhibitory synapses using endogenous proteins (e.g., intrabodies - FingR or nanobodies) should be included.

      This important point was also raised by other Reviewers. We have now significantly expanded the Discussion section, including discussion of this point.

      (3) The activity state of a neuron and/or a synapse might alter the dynamic properties (formation, maintenance, and/or elimination). A discussion on whether the overexpression of Homer1 and/or gephyrin might alter synapse/neuron activity would provide greater interpretability of the results. A discussion of the potential limitations and benefits of the reporter and TKIT approaches would be beneficial.

      We agree and have added discussion of these points to the Discussion section.

      (4) A description and interpretation of the computational approach to calculate particle tracking would be helpful. I found that particle tracking figures, while elegant, are difficult to interpret.

      As discussed in more detail below, we have generated drift correction and particle tracking approaches for the revised manuscript. We now elaborate on these new approaches in the paper.

      We thank the Reviewer again for their very helpful input and suggestions.

      Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      We thank the Reviewer for their positive assessment of our study.

      However, a number of aspects remain to be addressed in order for the study to support the claims made by the authors. First, the novelty aspect of the development of the fluorescently tagged synaptic proteins is unclear, since reporters of this nature are in routine use in many labs. Second, the analysis of the acquired images often seems incomplete, with only example images but no quantification shown, or the distinction between spatial and temporal dynamics appearing unclear. Third, given this incomplete analysis, the interpretations of the authors are not always convincingly supported by the data presented. In conclusion, substantial improvements are required to render the main messages of the study clear and compelling.

      We agree and have incorporated all of the Reviewer’s suggestions in the revised manuscript (please see below).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This is an interesting study. This reviewer has the following questions/comments for the authors:

      (1) Please provide evidence that the gRNAs targeting each gene of synaptic protein have no offtarget effects.

      We now include analysis of off-target effects for the TKIT tools (new Figure S6).

      (2) While structural E/I balance is shown, functional electrophysiological validation (e.g., mEPSC/mIPSC ratios) is absent. It is interesting to know whether the balanced functional structural changes translate to functional?

      We thank the Reviewer for this insightful suggestion and now include these recordings in the revised paper (new Figure 8).

      (3) In lines 217-218, please define thresholds for "stable" vs. "dynamic" puncta (e.g., temporal and spatial criteria).

      We more clearly define our categorization parameters (e.g. new Figure 2).

      (4) In Figure 5B: The low co-localization between endogenously tagged Bassoon and antibodystained Bassoon is likely due to the low TKIT efficiency. Quite a few HA-tagged Basson signals are insensitive to Basson-antibody. The authors are suggested to explain those.

      We thank the Reviewer for identifying this and add discussion to the Results section.

      (5) For the data analysis. If each n represents an independent neuronal culture, should the authors are suggested to provide the number of neurons/dendrites analyzed for each independent culture?

      We have added these important details to the manuscript.

      (6) Regarding the title, the author used the term "coordinated dynamics". This reviewer finds it is a bit over-claim because the stable ratios of the number of excitatory synapses and inhibitory synapses are likely an association, not actively "coordinated". I suggest that the authors rephrase this.

      We agree that we cannot argue that excitatory and inhibitory synapses are causally coordinated in our current study. Their levels are likely associated by either association or direct coupling, which we now discuss further in the first paragraph of the Discussion. We have rephrased the title accordingly.

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions that I think will improve the manuscript:

      (1) Please define Syn1/2 on line 129.

      We have defined this in the revised paper.

      (2) For Figures 2B, C, and 4B, C: are the puncta in panel C from the dendrites in panels B? If so, it would be helpful to identify the ROIs selected in panels C.

      We now include this in new Figure 2.

      (3) For the particle tracking figures, while the ability to track all synaptic puncta is very impressive, it is sometimes difficult to clearly track the lifespan of a synaptic puncta from the current figures. I believe that it would be helpful if the authors selected specific examples of synapses formed, maintained, and eliminated.

      We agree and now include more examples.

      (4) I believe that more detail about the computational approach and analysis for the particle tracking (Figs 2E and 4E) would help the interpretability of the figure.

      This important point was also raised by the other Reviewers. We generated custom tools during the revision that significantly expand the capabilities of our tracking approaches and more clearly describe them in the revised manuscript.

      (5) Similar to the rigorous gephyrin TKIT analysis (Fig. 6), did the authors perform a similar analysis for Homer1c TKIT? This might be valuable to confirm that overexpression of the Homer1 reporter does not indirectly alter synapse dynamics.

      We attempted to perform live imaging of mClover3 TKIT-tagged endogenous Homer1 but encountered low signal/noise with live imaging. We now add discussion that optimization of more robust tags (e.g. StayGold, HaloTag) will likely be necessary for live imaging of different target proteins.

      (6) The tools developed by Garbett et al. have the potential to be broadly utilized in the field to provide new insight into the coordination of excitatory and inhibitory synapses. It would thus be helpful for the authors to include a discussion about the strengths and limitations of the reporter and TKIT methods relative to other approaches used to live image synapses (e.g., intrabodies (FingR and nanobodies)).

      We have now significantly expanded the Discussion to include these important points.

      (7) In the discussion, can the authors elaborate on whether it is experimentally feasible to apply their TKIT labeling of gephyrin and Homer1c in the same neuron to assess the endogenous excitatory and inhibitory synapse dynamics from the same neuron?

      We have added discussion of this point and also proof-of-concept data supporting tagging of two postsynaptic targets within the same neuron (new Figure S5D).

      Reviewer #3 (Recommendations for the authors):

      (1) While the new tools described in the current manuscript can undoubtedly be used for the described purposes, the novelty of these tools is unclear to me. Viral vectors expressing fluorescently tagged versions of Homer1, synaptobrevin, and gephyrin are commercially available, e.g., via Addgene, and they are in routine use in many labs. CRISPR-mediated strategies for this purpose have also been previously reported (e.g., Willems et al. 2020, PLOS Biology; Fang et al. 2021, eLife). It is not clear to me how the tools reported here present a significant improvement over existing resources, other than that they use different fluorescent tags. If this aspect is a central part of the current manuscript, it should be expanded on in the discussion, including a direct comparison with available tools to highlight the novel aspects.

      We agree and have significantly expanded the Discussion to include these important points. Also, rather than argue that our tools are superior to pre-existing approaches, we adjust the text to argue that our tools and analytical approaches have been designed and optimized for the purposes we apply them to.

      (2) In addition to generating new tagged constructs, the authors also state that they have developed new imaging and analysis strategies to facilitate long-term assessment of synaptic dynamics. However, in many figures, they present only sample images, with little quantification to allow assessment of the wider relevance of the imaged synapses. For example, in Figures 2C and 4C, they present one example each of, e.g., a stable, nascent, transient, or eliminated synapse. However, they do not provide any quantification on how frequently any of these events occur, or whether they can be reliably quantified at all. These quantifications (i.e., percentage of each event type across a large population of synapses) would be necessary and should be added to demonstrate that this tool can be used for more than single example images.

      We have generated custom-made drift correction and particle tracking approaches for the revised manuscript. Based on the reviewer’s suggestion, we have quantified the relative frequencies of stable, nascent, transient, and eliminated synapses (Fig 2B-G, Fig3A-F, Fig 5A-F, Fig 7B-C). These metrics greatly enhance the biological interpretation of our results. We have also added a supplemental movie with an example image with corresponding categorized tracks for each puncta type (Movie S3)

      (3) The authors do present an automated visual representation of spatial track length across the neuron, e.g., in Figure 2E and 4E, although this is also not quantified. Moreover, the track lengths appear surprisingly short, despite the authors' claims that their analyses 'highlight the dynamic nature of excitatory synapses over these timescales'. It is not clear to me whether these short tracks are more than just jitter, either in the synapses themselves or in the images due to technical limitations. E.g., in panel 2E, I see very few examples in which the track is not simply centered around one point, but actually expands over a distance. Quantification of the distance between start and end points of the tracks would be important to support the claim that these synapses are dynamic in terms of spatial translocation (if that is what the authors meant). Or if the 'dynamic nature' of the synapses referred to temporal dynamics, it is unclear to me how this information can be gained from the represented tracks.

      We thank the reviewer for these excellent points. To accurately access spatial motion, we drift-corrected our images with a custom correction algorithm to eliminate stage or microscope drift as a source of contaminating motion (See Methods, Movie S2), in addition to collecting time-lapse imaging with Nikon perfect focus. We noticed heterogeneity in our cultures such that some areas contained very mobile neurites, while other remained stationary (Fig. S1). We binned movies into either moving or still neurites and assessed spatial metrics as suggested (Fig. S1A). Consistent with our binning, puncta on moving neurites showed larger net displacement (distance between start and end points), but puncta on still neurites also showed ~1 µm net displacement (Fig. S1D). We also quantified puncta speed and found that puncta on moving neurites generally moved faster (Fig. S1C). We appreciate the reviewer’s insight that track length were surprisingly short, and after employing our drift correction and revised tracking methods, we now see substantially longer track lengths (Fig 2E, Fig 3C & F, Fig S2B & C). We additionally see a large fraction of tracks that persist throughout the imaging session (Fig 2E, Fig S2B & C).

      (4) In Figure 3, the authors now quantify track length, but in this case in the unit 'minutes', from which I would interpret that this is now meant to assess the temporal dynamics rather than the spatial dynamics. The lack of a clear distinction between spatial dynamics and temporal dynamics is very confusing to me, since these are entirely independent measures. 'Track length' to me indicates spatial dynamics, and I would expect the units to be a measure of distance. 'Track duration', which the authors also use in some places, but inconsistently as far as I can tell, makes sense to me for the assessment of temporal dynamics, with the units being a measure of time. I would strongly recommend being very clear about this distinction, since the current representation of the data is very difficult to follow and interpret.

      In addition to new spatial metrics, we have clarified in the text when we are referring to spatial dynamics (distance) versus temporal dynamics (time). As suggested, we use duration when referring to time, and speed or distance when referring to spatial metrics.

      (5) The images from the newly generated CRISPR-based tags in Figures 5-7 are striking and very compelling - these will be very useful tools. However, here too, it seems that the interpretation of the data does not really match the results. All quantification indicates that there is very little change in synapse density or other assessed parameters over the time course of the imaging, and yet the authors emphasize the dynamic nature of visualized synapses. More compelling quantification would be needed to support this claim.

      We have quantified spatial and temporal metrics for live neuron culture imaging for all tools developed including CRISPR-based tags (Figure 7).

      (6) The discussion is extremely short and provides almost no integration of the results of the study into the framework of existing knowledge. Instead, it focuses almost exclusively on unanswered questions and future perspectives, which are also important, but not helpful in interpreting the findings from the current study. The latter aspects should be added to provide essential context for the current findings.

      We agree and have added additional discussion of our current findings to help contextualize their significance.

      We thank the Reviewers again for their positive feedback and insightful input, which has undoubtedly strengthened our study.

    1. eLife Assessment

      This important study shows that regions of the human auditory cortex that respond strongly to human voices are also sensitive to vocalizations from closely related primate species. The evidence is convincing and methodologically strong. The work offers significant insight into the evolutionary continuity of voice processing and would be of interest to researchers studying auditory processing and evolutionary neuroscience in general.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance u

      Comments on revised version.

      I thank the authors for having carefully considered and implemented my remarks on the first version.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding-that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations-suggests a potential neural sensitivity to calls form phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weakness:

      The authors only tested vocalizations from three non-human primate species other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      Comments on revised version.

      I have no further comments.

    4. Reviewer #3 (Public review):

      Summary:

      Using fMRI, the authors demonstrate that human temporal voice areas (TVA) respond not only to human vocalizations but also to those of other primates, particularly chimpanzee calls, which share acoustic features with human voices. These findings provide compelling evidence for cross-species vocal processing in the human auditory system and carry important theoretical implications for understanding the evolutionary underpinnings of speech perception.

      Strengths:

      The study offers a valuable comparative design, rigorous acoustic and phylogenetic modeling, and consistent evidence that bilateral anterior TVA regions respond more strongly to chimpanzee vocalizations than to other species' calls. The inclusion of both great apes and monkeys provides a rare cross-species perspective.

      Weaknesses:

      Minor limitations include the acoustic-phylogenetic confound (which the authors partially address with additional analyses), the lack of non-vocal controls to establish true selectivity.

      Overall, the methods, data, and analyses broadly support the claims, with only minor weaknesses that do not undermine the main conclusions. The findings are valuable for the subfield of auditory neuroscience and comparative cognition, with solid evidence supporting the primary claims.

      Comments on revised version.

      After revision, this work has shown great improvement in data analysis, figure organization, and writing. I have no further suggestions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance understanding of the neural bases of voice perception and the evolutionary roots of voice sensitivity in the human brain.

      Weaknesses:

      (1) Acoustic-phylogenetic confound: The design does not fully disentangle acoustic similarity from phylogenetic proximity, as species co-vary along both dimensions. A promising way to address this would be to include an additional model focusing on the acoustic features that specifically differentiate bonobo from chimpanzee calls, which share equal phylogenetic distance to humans.

      (2) Selectivity vs. sensitivity: Without non-vocal control sounds, the study cannot determine whether TVA responses reflect true selectivity for primate vocalizations or general auditory sensitivity.

      (3) Task demands: The use of an active categorization task may engage additional cognitive processes beyond auditory perception; a passive listening condition would help clarify the contribution of attention and task performance.

      (4) Figures and presentation: Some results are partially redundant; keeping only the most representative model figure in the main text and moving others to the Supplementary Material would improve clarity.

      We thank the reviewer for contributing to the improvement of the present study and for the extremely constructive criticism. Concerning the identified weaknesses of our work, we provide here some general answers while the detailed review (below) addresses point-by-point the reviews in high detail.

      (1) We totally agree that acoustics and phylogeny cannot be disentangled in our study, which is a limitation. We now provide the suggested analysis on the acoustic specificities of chimpanzee and bonobo calls.

      (2) This point on selectivity vs. specificity is indeed crucial, and we now provide a more careful viewpoint and phrasing on this aspect, since our study can only provide partial arguments for this important distinction.

      (3) Task demand following species categorization might rightfully yield to the engagement of distinct brain network compared to merely listening to the stimuli. We discuss this aspect and put forward the argument that, while we cannot control for this aspect, our attentional control study performed by an independent sample, N=28 provides clear evidence that no species triggered an attention bias. In other words, task demand might play a role, but at least in the study we know that attentional resources were not biased towards one species in particular since no effects were observed.

      (4) We agree that results were not articulated in a clear fashion and that figures were redundant. We addressed this aspect and regrouped the figures where appropriate while we include the rest in the supplementary material now.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding - that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations - suggests a potential neural sensitivity to calls from phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weaknesses:

      The interpretation of the findings in this paper regarding the evolutionary continuity of voice processing lacks sufficient evidence. A simple explanation is that the observed effects can be attributed to the similarity in low-level acoustic features, rather than effects specific to phylogenetically close species. The authors only tested vocalizations from three non-human primate species, other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      We want to thank the reviewer for the constructive criticism and for evaluating the manuscript.

      Concerning the principal weakness highlighted, we provide new analyses behavioral, acoustics, model-based fMRI that improve our understanding of the influence of both phylogeny and bioacoustics in our data. We argue that the explanation proposed by the reviewer cannot explain our results, as also observed in several other research from us and others. We discuss this aspect and emphasize that including stimuli from more species would greatly improve the understanding of phylogeny and bioacoustics in this context.

      Reviewer #3 (Public review):

      Summary:

      Ceravolo et al. employed functional magnetic resonance imaging (fMRI) to examine how the temporal voice areas (TVA) in the human brain respond to vocalizations from different nonhuman primate species. Their findings reveal that the human TVA is not only responsible for human vocalizations but also exhibits sensitivity to the vocalizations of other primates, particularly chimpanzee vocalizations sharing acoustic similarities with human voices, which offers compelling evidence for cross-species vocal processing in the human auditory system. Overall, the study presents intellectually stimulating hypotheses and demonstrates methodological originality. However, the current findings are not yet solid enough to fully support the proposed claims, and the presentation could be enhanced for clarity and impact.

      Strengths:

      The study presents intellectually stimulating hypotheses and demonstrates methodological originality.

      Weaknesses:

      (1) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy.

      We thank the reviewer for evaluating our manuscript and for the constructive criticism as well as the many suggestions. Concerning the weaknesses of the study, we provide here some quick answers while more detailed responses can be found below.

      (1) We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator.

      (2) We totally agree that figure redundancy was a problem and we now reduced confusion by combining congruent aspects while pushing other results to the supplementary material.

      Recommendations for the authors:

      Reviewing Editor Comments:

      With additional analyses and discussions, the work has the potential to offer important insight into the evolutionary continuity of voice processing.

      We thank the Reviewing Editor for this additional motivation and for offering us the possibility to revise our manuscript. We will now provide our point-by-point reviewing, referring to manuscript modifications by section and/or line number(s). All modifications are also highlighted in light grey in the text.

      Reviewer #1 (Recommendations for the authors):

      The manuscript is clearly written and addresses an important comparative question about the specificity of human TVA responses. The acoustic analyses are well designed, and the imaging work is careful and thorough. However, several conceptual and methodological issues need clarification or tempering of claims, particularly regarding (i) the distinction between sensitivity and selectivity, (ii) the confounding of acoustic and phylogenetic factors, and (iii) the interpretation of "chimpanzee-specific" TVA activity.

      (1) Introduction

      Line 48: cite more recent infant EEG evidence for early voice sensitivity (Calce, Curr Biol).

      The reference and explanation were added, lines 46-48.

      Line 53: mention recent data on voice processing in marmosets (Jafari, Cell Rep; Dureux, Curr Biol).

      We added the references and the mention of these interesting studies on common marmosets, lines 53-54.

      Line 59: Fecteau et al. (2004) already explored cross-species selectivity; please integrate and discuss.

      We now mention here the work from Fecteau and colleagues and its relevance, see lines 57-59.

      Line 70: clarify that in [27] (Bodin et al., 2021) human TVA responded similarly to human nonverbal vocalizations and macaque coos, likely due to acoustic similarity.

      We added this important aspect, thank you for this precision. See lines 71-72.

      Clarify why an active species-categorization task was chosen instead of passive listening, which is standard in TVA research. Were participants familiarized with stimuli beforehand?

      We added a sentence on this aspect, but basically to summarize it here: we wanted to be able to test human recognition of nonhuman primate species’ calls. From the start, we wanted to test the frontal mechanisms related to decision-based processes of humans when categorizing non-human primate calls hence the 2023 article we published. See lines 75-77 and we also added information on familiarization to the stimuli in the Methods, lines 679-682.

      The 16 acoustic features mentioned should be briefly defined earlier, as they are central.

      We feel like describing 16 acoustic parameters in the introduction would be heavy on the reader, so we instead added a reference to the supplementary table (Table S1) in which these are named and described. See line 80.

      Explain why only chimpanzees and bonobos were selected among the great apes, and discuss the value of including both, given their equal phylogenetic proximity but largely dissimilar acoustics.

      The stimuli were obtained by Thibaud Gruber and his team and through collaborations with Katie Slocombe and Zanna Clay. Unfortunately, at the time we could only use chimpanzee and bonobo calls for the great apes. Therefore, it was mainly a material constraint rather than a deliberate choice to exclude other great apes. We now discuss this aspect and present the absence of other great apes as a limitation (lines 587-591).

      Rephrase references to "recruitment" of TVA - this term implies general activation, while the key question concerns selectivity (stronger responses to voices vs. non-vocal controls).

      We rephrased throughout the manuscript, thank you for this suggestion.

      The hypothesis section should more clearly separate the acoustic and phylogenetic predictions, and clarify which earlier data motivate each.

      We now explicitly categorize the hypotheses according to either Bioacoustics or Phylogeny to clarify. We also added references motivating each hypothesis. See lines 114-120.

      (2) Methods

      Clarify whether stimuli were RMS-normalized or otherwise balanced for energy (line 128).

      Sound pressure level was kept constant but the stimuli were not normalized, specifically to avoid a negative impact on their naturality. We added a sentence (lines 131-132) including a reference on this aspect.

      The task design could benefit from reporting accuracy in addition to reaction times for the 4AFC species classification task.

      We agree this aspect was missing. We now report accuracy data (controlled for reaction times and acoustics of Model 3) for the species categorization task (lines 147-165; Fig.1B), and in the Methods (lines 769-786). The fitted regression values of this analysis are also used for a new fMRI model (Model 4), to uncover within-TVA correlates of the probability of correct species categorization (lines 309-325; Fig.4).

      Please note that previously, the behavioral data of the species categorization task were completely absent (N=23), and the reaction times data previously part of Fig.1 were for the species attentional bias task (independent sample of N=28). Since this aspect was not clear at all (same remark by all reviewers—apologies for that), we now include a clear separation in Fig.1, with newly added panels D & E part of a distinct figure area named: “Control task: Testing for Species attentional bias (N=28)”. Panel D illustrates the control task paradigm (each species as exogenous cue; “dot-probe” paradigm) while panel E shows the results (target sine wave tone or “bip” detection reaction times), showing that no species triggered more attentional capture than the others (Species effect non-significant).

      The acoustic parameters used in Models 2 and 3 should be explicitly listed in the Methods (even if already published elsewhere).

      In addition to their description in Table S1, we now include the 16 acoustic parameters used to calculate acoustic distance between the species in the Methods, see lines 828-844.

      Consider simplifying the presentation of the three models: a figure summarizing their relationships would help.

      We now include only one figure (Fig.2) for Model 3, and we pushed model 1&2 to the supplementary material. We also simplified Fig.3 for a clearer view of the overlaps between the 3 models within the TVA.

      The description of “systematic and thorough control of phylogeny” (line 119) is overstated, given that only three nonhuman species were included.

      We agree with the reviewer and we suppressed both “systematic” and “thorough” from the sentence.

      Provide rationale for not including a nonvocal control category (e.g., scrambled vocalizations or environmental sounds) to assess TVA selectivity.

      The main objective of the study was to uncover whether human participants could recognize the vocalizations from nonhuman primates—from both great apes and monkeys—as compared to the human voice. We therefore did not include nonvocal or noise stimuli. We added this point as a limitation in the Discussion (lines 593-596 and 609-611).

      Even though we did not include such stimuli for the reason mentioned above, the delineation of subtypes of nonvocal material within the TVA of our participants (Fig.2) are, in our opinion, clarifying the message: chimpanzee-selective activations are fully within ‘voice vs. animal’ and ‘voice vs. nature’ TVA subareas, while it is not the case in ‘voice vs. music’ and ‘voice vs. noise’ TVA subareas.

      Clarify if participants were trained or had a practice session to recognize the four species before scanning.

      The participants were indeed trained on 3 stimuli per species before entering the MRI scanner. These stimuli were discarded from the species categorization task. We added a sentence about this aspect, see lines 131-132.

      Specify what is meant by "no good or bad response" in the attentional control task (line 724).

      We suppressed this wording as it was highly confusing.

      (3) Results

      Behavioral accuracy should be reported to complement reaction times.

      We now added behavioral data for the species categorization task as well as the neural correlates of accurate species categorization. See our previous response above (‘‘‘).

      Figures 2-4 largely overlap; consider merging or simplifying to reduce redundancy.

      We agree and this point was raised by the other reviewers as well. Task-based results are now presented only for Model 3 as Fig.2, while Fig.3 (previously Fig.5) summarizes the overlap between the three models. Figures for Models 1 & 2, previously labelled Fig.3 and Fig.4, were moved to the supplementary material.

      Figure 2: Please indicate more clearly where "chimp-selective" areas are located (perhaps with zooms).

      We agree, we now modified Fig.2 with zoomed-in panels and a clearer outline of chimp-selective areas (solid blue outline). This outline is also referenced in the text (lines 236-237).

      Correction for multiple contrasts: With many pairwise tests, adjustments (Bonferroni or FDR) should be mentioned explicitly.

      We now specify ‘FDR correction at the voxel level’ at the beginning of the Results section (lines 195-198) as well as in each figure.

      Replace "specific to chimpanzee" with "selective for chimpanzee" to avoid implying exclusivity.

      We made the suggested replacement throughout the manuscript.

      Discuss whether the small macaque-related clusters might simply reflect acoustic overlap rather than true category selectivity.

      We added a section on this important aspect, including results that support the role of mid-STG/STS regions for more noise-like stimuli, including the use of macaque coos. See lines 450-461.

      (4) Discussion

      The discussion overstates claims of "chimpanzee-selectivity" in TVA. The evidence shows relative preference, not absolute selectivity.

      We now specify from the start of the Discussion that we are not interpreting the results as absolute selectivity but rather as more relative preference, see lines 371-373.

      The authors repeatedly conflate acoustic and phylogenetic factors; this should be explicitly acknowledged as a limitation.

      We agree, and we completed the limitations section already dedicated to this aspect by a more explicit account of the confound, see lines 609-611.

      Clarify what is meant by "recruitment" and "selectivity" (lines 411-419, 577). TVA activity often reflects enhanced responses to voices compared to non-vocal sounds, not exclusive activation.

      We clarified this wording in the Discussion (lines 377-378) and replaced another instance by “activated the […]” to make it clearer what we imply, namely enhanced activity triggered by chimpanzee calls within human TVA.

      The lack of non-vocal control conditions should be discussed as a major interpretive limitation.

      We added this point as a limitation in the Discussion (lines 593-596).

      The statement that "chimpanzee-selective activity" arose in humans who have never been exposed to chimp calls (line 450) invites evolutionary speculation but should be more cautiously phrased.

      We agree, and we rephrased by: “[…] with chimpanzee calls triggering responses in the anterior STG/TVA of our human participants […]”. See lines 432-433.

      The comparison to recent macaque data (Giamundo et al., 2024 PNAS) is crucial: these findings of human-voice-selective neurons in macaques directly parallel the present human-chimp result.

      We agree with the reviewer, and we are hopeful to read similar results for other apes/great apes in the future.

      Reviewer #2 (Recommendations for the authors):

      (1) The primate vocalizations used in this study were recorded in diverse social and emotional contexts, which may have contributed to the observed differences in TVA activation. Since the temporal voice areas are known to be sensitive to affective and socially relevant cues, these contextual differences could confound the interpretation of species-specific neural responses. Therefore, I suggest that the authors conduct a post-hoc analysis to quantify and compare the affective valence, arousal levels, and social contexts associated with each stimulus set.

      We agree that the TVA are sensitive to social—or socially relevant—cues, motivating the very thorough work of the expert reserve personnel on-site to accurately categorize the calls according to the very specific context they were produced in. If the reviewer meant presenting these stimuli to non-expert participants and asking them to categorize the context or valence, we think it would make no sense since the ratings would be completely below chance level and therefore uninformative. The newly added behavior—and model-based fmri—data include this crucial point, a factor that we named ‘Context’ in our analyses. In fact, for each species’ 18 stimuli, we control for agonistic and affiliative production context—split evenly, per species. Also, computing an additional posthoc analysis by splitting the stimuli according to Context would result in too few trials to get sensible and reliable fMRI results.

      That being said, our study targets this specific aspect by extracting the acoustic features that characterize our stimulus set the best, across context-species-valence-arousal, which is exactly what we want. Through the three types of modeling we used—from more simplistic to more elaborate the results converge only for one species: chimpanzee calls.

      We think the addition of behavioral data, model-based fMRI data, and the specific analysis on acoustic differences between chimpanzee and bonobo calls strengthens the message and the validity of our findings.

      (2) Although the author mentioned that the behavioral effects triggered by these vocalizations have been reported previously, the behavioral responses of the participants in the current study are also crucial for our understanding of the results. If the MRI data can be combined with the participants' behavioral responses for comprehensive analysis, the conclusions of this study will be more compelling.

      We agree with the reviewer, and we added the behavioral data—controlling for reaction times, production context and acoustics of interest—and we also included a model-based fMRI modeling of the probability of correct species categorization as Model 4, Fig.4. See, respectively: lines 147-165, Fig.1B; Methods, lines 769-786; Neuroimaging results, lines 309-325.

      (3) I am still not convinced that phylogenetic proximity drives the observed neural selectivity. While chimpanzee vocalizations do elicit stronger responses in anterior STG, the claim that this reflects evolutionary relatedness lacks evidence. If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree with the reviewer that generalizing our results in terms of phylogenetic proximity alone is not a viable option. Including many more primate species including other great apes would be necessary, and we mention this crucial aspect in the limitations section. We also insist in the Discussion on the interdependence between phylogeny and acoustics in our data, since: 1) we cannot fully disentangle these factors here, 2) we cannot attribute our results to either one or the other. See lines 387-390, 410-411, 473-477, 587-591.

      If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree, and nobody could disagree: if an auditory object is extremely similar to the human voice in terms of acoustics, it would therefore potentially activate the TVA. This is exactly our message: in the natural ‘auditory world’, the calls from chimpanzees seem to be among the very few animal auditory signals that are sufficiently close, acoustically, to the human voice and therefore trigger TVA activity. They also happen to be the calls from a species which is phylogenetically the closest to humans with minimal differences with other great apes. Our results are in that sense very aligned with work from the laboratory of Pascal Belin, namely on ‘voice patches’ in the primate brain located in the (anterior) TVA, cited in our manuscript.

      We therefore think our interpretation does not exclude that in the near future, similar results within the TVA could be observed for other auditory objects, and if animal, from a species potentially much more distant phylogenetically or from vocal signals of other great apes.

      We added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as scrambled or spectrum shifted per-species stimuli, which would have made the interpretation clearer identical acoustics but alteration/destruction of the species auditory object. See lines 593-596 and 609-611.

      Reviewer #3 (Recommendations for the authors):

      While the manuscript presents intriguing results, several concerns are raised for further consideration, detailed below.

      We thank the reviewer for evaluating the manuscript and for the constructive criticism and suggestions.

      Major concerns:

      (1) This study claims that bilateral anterior superior temporal gyrus (aSTG) in humans can be specifically activated by chimpanzee vocalizations rather than all other primate species after regressing out relevant acoustic parameters using three distinct analyses. I am wondering if a control stimulus (e.g., scrambled chimpanzee vocalizations) were presented, would the activation patterns in these same temporal voice areas (TVA) exhibit significant differences compared to the natural chimpanzee vocalizations?

      We completely agree with the reviewer, and this point was also raised by the other reviewers. We therefore added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as per-species scrambled or spectrum shifted stimuli, which would have made the interpretation clearer—identical acoustics but alteration/destruction of the species auditory object. See lines 609-611.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy. E.g:

      Figure 1C is the same as Figure S1. In addition, Figure 1C lacks a figure legend and descriptive label.

      The scatter plots in Figures 2D, 2H, 3D, 3H, and 4D, 4H are same as those in Figures S2, S3, and S4. However, some of these duplicate plots even have inconsistent axis labels.

      In several panels, the main figures appear to be summaries derived from the supplementary figures. The authors should organize these figures well to eliminate redundancy.

      Please double-check all the figures to make sure of accuracy.

      We agree that the figures were badly organized and were too crowded and redundant. We now suppressed the redundancy between Fig.1 and Fig.S1, and we reduced fMRI results to one figure for statistical Model 3 while the other models are in the supplementary data—we also justify this decision in the text by highlighting that model 3 is the most elaborate and sensitive one. Fig.3 (previously ‘Fig.5’) shows the overlaps between models and was simplified and clarified as well.

      (3) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task. It is possible that processing vocalizations from certain species requires more cognitive effort or induces higher decision uncertainty. Could the observed neural effects be confounded by the decision-making process itself?

      We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator. We now display these results in Fig.4 and we introduce the motivation factor for including a categorization task rather than more traditional passive listening (lines 75-77), as well as limitations, lines 595-596.

      (4) One interesting attempt of this study is to dissociate biologically salient information in animal vocalizations from their low-level acoustic properties. This presents a fundamental conceptual challenge: how to rigorously disentangle a vocalization's species-specific attributes from its inherent acoustic correlates. More precisely, what essential biological information persists in a species' vocal signal after statistically accounting for all quantifiable acoustic features? I recommend that the authors address it in the discussion.

      We thank the reviewer for this very important comment, and for suggesting we discuss it in the manuscript. We completely agree: we cannot fully orthogonalize species and acoustics, and this aspect relates also more broadly to cognitive and affective neuroscience studies involving vocal material. Namely: “What is an auditory object without acoustics?”

      We included a full paragraph on this aspect, see Discussion, lines 570-584.

      (5) If a brain region, such as TVA, is responsive to both acoustic parameters and biological meanings of animal vocalizations, the method used in this study might be inadequate by setting covariates to zero. It is possible that species information is embedded within a specific acoustic pattern. The current modeling approach may not capture such complex information and could potentially introduce bias when estimating the species effect. I recommend that the authors address this issue in the discussion.

      We thank the reviewer for this point once again, we addressed it in the Discussion, lines 581-584, and also in the section dedicated to study limitations, lines 609-613.

      (6) In the discussion, non-human primate vocalizations are "unreadable" to humans. If this is the case, what is the fundamental perceptual difference between these vocalizations and those from the other animal species? An alternative and highly plausible explanation for the findings is the differential familiarity of the participants with the various species, driven by media exposure (e.g., documentaries) or zoo visits and interactions. The authors need to provide a stronger justification for their control stimuli and directly address, either through discussion or additional analysis, how the factor of familiarity might explain their results better than the proposed "evolutionary distance" hypothesis.

      We now discuss this important aspect, see lines 560-569.

      We thought about doing additional analyses on this aspect but we concluded that we did not have any reliable indicators of familiarity for our participants, and additionally they were all recruited for being ‘unfamiliar’ with great apes or old-world monkeys’ vocalized communication.

      Also, frequent mismatches in the media between images of apes and the associated vocal signals (for instance, the depiction of a chimpanzee but with background audio of macaque coos) are not helping this cause.

      Minor:

      (1) No figure legend and result description for Figure 1.

      Figure 1 has a legend, maybe it was cut out during the uploading process, but it is present and verified now.

      (2) In the main text, three statistical models were referenced. Was the data used in each subsequent statistical model derived from the processed data of the preceding model? Please clearly explain this in the main text.

      We now specify this aspect in the Methods and the Results section to clarify that each model is independent from the others (lines 964-966 and 189-191, respectively).

      (3) In Figure 5, the two dashed lines representing Model 1 and Model 2 are confusing for readers.

      We modified the figure (now Fig.3) and simplified it by removing some outlines and clarifying the colors, therefore improving readability.

      (4) Lack of reaction times in the species categorization task.

      We clarified behavioral data, including the results for the species categorization task and for the control, exogenous cueing task, see modified Fig.1 and behavioral results section of the Results.

      (5) Figures 2, 3, 4, 5, Please keep the font size of the figure title consistent.

      Figure title font size were uniformized.

      (6) Line 201, Line 224, and so on, (EFG) → (E, F, G).

      We modified this aspect in every figure legend, including the supplementary material.

    1. eLife Assessment

      Argunşah et al. investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in shaping the responses to single vs multiple whiskers. Based on the observation of a higher density of SST+ interneurons in the septa, the authors investigate the hypothesis that Elfn1-dependent short-term plasticity shapes these responses. This important study is, however, supported by incomplete evidence; factors restricting the strength of evidence are the limited spatial resolution of the multi-unit activity, as well as the lack of a mechanistic explanation. This provocative and intellectually stimulating hypothesis provides a contribution to work on how different cell types shape cortical representation.

    2. Reviewer #1 (Public review):

      Summary:

      Argunşah et al. describe and investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in the whisker-related primary somatosensory cortex (S1). Upon repeated stimulation, the authors report that the response ratio between multi- and single-whisker stimulation increases in layer (L) 4 neurons of the septal domain, while remaining constant in barrel L4 neurons. The authors attribute this divergence to differences in short-term synaptic plasticity, particularly within somatostatin-expressing (SST⁺) interneurons. This interpretation is supported by 1) the increased density of SST+ neurons in L4 of the septa compared to barrel domain, 2) the stronger response of (L2/3) SST+ neurons to repeated multi- vs single-whisker stimulation and 3) the reduced functional difference in single- versus multi-whisker response ratios across barrel and septal domains in Elfn1 KO mice, which lack a synaptic protein that confers characteristic short-term plasticity, notably in SST+ neurons. Consistently, a decoder trained on WT data fails to generalize to Elfn1 KO responses. Finally, the authors report a relative enrichment of S2- and M1-projecting cell densities in L4 of the septal domain compared to the barrel domain, suggesting that septal and barrel circuits may differentially route information about single vs multi-whisker stimulation downstream of S1.

      Strengths:

      This paper describes and aims to study a circuit underlying differential response between barrel columns and septal domains of the primary somatosensory cortex. This work supports the view these two domains contribute distinctly to the processing single versus multi-whisker inputs and highlight the role of SST+ neuron and their short-term plasticity. Together, this study suggests that the barrel cortex multiplexes whisker-derived sensory information across its domains, enabling parallel processing within S1.

      Weaknesses:

      Although the divergence in responses to repeated single- versus multi-whisker stimulation between barrel and septal domains is consistent with a role for SST⁺ neuron short-term plasticity, the evidence presented does not conclusively demonstrate that this mechanism is the critical driver of the difference. The lack of targeted recordings and manipulations limits the strength of this conclusion: SST⁺ neuron activity is not measured in L4, nor is it assessed in a domain-specific manner. The Elfn1 knockout manipulation does not appear to selectively affect either stimulus condition, domain or interneuron subtype. Finally, all experiments were performed under anesthesia, which raises concerns about how well the reported dynamics generalize to awake cortical processing.

    3. Reviewer #3 (Public review):

      Summary:

      This study investigates the functional differences between barrel and septal columns in the mouse somatosensory cortex, focusing on how local inhibitory dynamics (particularly involving SST⁺ interneurons) may mediate temporal integration of multi-whisker (MW) stimuli in septa. Using a combination of in vivo multi-unit recordings, calcium imaging, and anatomical tracing, the authors propose a model in which Elfn1-dependent synaptic facilitation onto SST⁺ interneurons contribute to the distinct sensory responses to MW input in barrels and septa, enabling functional segregation between these domains.

      Strengths:

      The study presents a thought-provoking and useful conceptual model for understanding sensory processing in the somatosensory cortex. While barrel columns have been widely studied, septal regions remain relatively understudied in mice. If septa indeed act as selective integrators of distributed sensory input, this would suggest a novel computational role for cortical microcircuits beyond the classical view focused on barrels. Although still hypothetical, the proposed model in which SST⁺ interneurons contribute to domain-specific sensory responses between barrel and septal domains is intriguing and opens new avenues for investigating inhibitory circuit mechanisms.

      Weaknesses:

      The primary limitation of this study lies in the spatial and cellular specificity of the recording techniques. The physiological data rely predominantly on unsorted multi-unit activity (MUA) recorded with low-channel-count silicon probes. Because MUA aggregates signals from multiple neurons over a radius of approximately 50-100 µm (comparable to or larger than the width of septal domains in mice), it remains difficult to confidently attribute the recorded activity exclusively to septal versus barrel populations. The authors have now addressed this concern more carefully by reframing their interpretation in terms of "septal-enriched" populations and by providing additional threshold-based analyses suggesting that the principal effects are more robust in Layer 4. These additions substantially improve the manuscript and support a more cautious interpretation of the findings. Nevertheless, the proposed Elfn1/SST⁺ mechanism remains supported primarily by indirect evidence. Although the calcium imaging data provide useful support for stimulus-dependent SST⁺ recruitment, these experiments were restricted to L2/3 interneurons and therefore do not directly test the Layer 4 circuit mechanism proposed to underlie the electrophysiological observations. Direct in vivo cell-type-specific recordings and manipulations in Layer 4 would ultimately be required to establish the proposed mechanism more conclusively.

      Comments on revised version.

      I have read the revised manuscript and overall, I think the authors have addressed my major concerns appropriately. I appreciate the substantially moderated interpretation of the findings and the additional analyses clarifying the limitations of the MUA recordings.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Argunşah et al. describe and investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in the whisker-related primary somatosensory cortex (S1). Upon repeated stimulation, the authors report that the response ratio between multi- and single-whisker stimulation increases in layer (L) 4 neurons of the septal domain, while remaining constant in barrel L4 neurons. The authors attribute this divergence to differences in short-term synaptic plasticity, particularly within somatostatin-expressing (SST<sup>+</sup>) interneurons. This interpretation is supported by

      (1) The increased density of SST+ neurons in L4 of the septa compared to barrel domain,

      (2) The stronger response of (L2/3) SST+ neurons to repeated multi- vs single-whisker stimulation and

      (3) the reduced functional difference in single- versus multi-whisker response ratios across barrel and septal domains in Elfn1 KO mice, which lack a synaptic protein that confers characteristic short-term plasticity, notably in SST+ neurons.

      Consistently, a decoder trained on WT data fails to generalize to Elfn1 KO responses. Finally, the authors report a relative enrichment of S2- and M1-projecting cell densities in L4 of the septal domain compared to the barrel domain, suggesting that septal and barrel circuits may differentially route information about single vs multi-whisker stimulation downstream of S1.

      Strengths:

      This paper describes and aims to study a circuit underlying differential response between barrel columns and septal domains of the primary somatosensory cortex. This work supports the view these two domains contribute distinctly to the processing single versus multi-whisker inputs and highlight the role of SST+ neuron and their short-term plasticity. Together, this study suggests that the barrel cortex multiplexes whisker-derived sensory information across its domains, enabling parallel processing within S1.

      Weaknesses:

      Although the divergence in responses to repeated single- versus multi-whisker stimulation between barrel and septal domains is consistent with a role for SST<sup>+</sup> neuron short-term plasticity, the evidence presented does not conclusively demonstrate that this mechanism is the critical driver of the difference. The lack of targeted recordings and manipulations limits the strength of this conclusion: SST<sup>+</sup> neuron activity is not measured in L4, nor is it assessed in a domain-specific manner. The Elfn1 knockout manipulation does not appear to selectively affect either stimulus condition, domain or interneuron subtype. Finally, all experiments were performed under anesthesia, which raises concerns about how well the reported dynamics generalize to awake cortical processing.

      We thank the reviewer for their careful reading of the manuscript and their balanced assessment of both its strengths and limitations. We acknowledge the reviewer’s concerns regarding the lack of direct, layer- and cell-type–specific recordings and manipulations of SST<sup>+</sup> interneurons, as well as the use of anesthesia. As noted in the Discussion, these factors limit the extent to which causal mechanisms can be established and the degree to which the reported dynamics can be generalized to awake cortical processing. For this reason, we intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, developmental, physiological, and genetic evidence, rather than as definitive proof. We believe this conceptual framing appropriately reflects the scope of the current data while highlighting clear directions for future work.

      Reviewer #2 (Public review):

      Summary:

      Argunsah and colleagues demonstrate that SST expressing interneurons are concentrated in the mouse septa and differentially respond to repetitive multi-whisker inputs. Identifying how a specific neuronal phenotype impacts responses is an advance.

      Strengths:

      (1) Careful physiological and imaging studies.

      (2) Novel result showing the role of SST+ neurons in shaping responses.

      (3) Good use of a knockout animal to further the main hypothesis.

      (4) Clear analytical techniques.

      Comments on revisions:

      The authors have effectively responded to my initial critiques - I have no further concerns.

      We thank the reviewer for their positive evaluation of our work and for recognizing the novelty of the findings, the careful physiological and imaging approaches, the use of the Elfn1 knockout model, and the clarity of the analytical framework. We are pleased that the reviewer has no further concerns and appreciates the contribution of this study to understanding the role of SST<sup>+</sup> interneurons in shaping sensory processing in the barrel cortex.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the functional differences between barrel and septal columns in the mouse somatosensory cortex, focusing on how local inhibitory dynamics (particularly involving SST<sup>+</sup> interneurons) may mediate temporal integration of multi- whisker (MW) stimuli in septa. Using a combination of in vivo multi-unit recordings, calcium imaging, and anatomical tracing, the authors propose a model in which Elfn1-dependent synaptic facilitation onto SST<sup>+</sup> interneurons contributes to the distinct sensory responses to MW input in barrels and septa, enabling functional segregation between these domains.

      Strengths:

      The study presents a thought-provoking and useful conceptual model for understanding sensory processing in the somatosensory cortex. While barrel columns have been widely studied, septal regions remain relatively understudied in mice. If septa indeed act as selective integrators of distributed sensory input, this would suggest a novel computational role for cortical microcircuits beyond the classical view focused on barrels. Although still hypothetical, the proposed model in which SST<sup>+</sup> interneurons contribute to domain-specific sensory responses between barrel and septal domains is intriguing and opens new avenues for investigating inhibitory circuit mechanisms.

      Weaknesses:

      The primary limitation of this study lies in the spatial and cellular specificity of the recording techniques. The physiological data rely predominantly on unsorted multi-unit activity (MUA) recorded with lowchannel-count silicon probes. Because MUA aggregates signals from multiple neurons over a radius of approximately 50-100 µm (often wider than the typical septal width in mice), this approach makes it difficult to confidently isolate activity originating strictly from within septal domains. The manuscript would benefit from additional analyses to validate the spatial specificity of these recordings, such as systematically varying spike detection thresholds to test the robustness of domain attribution, as suggested by the reviewer. Furthermore, although the authors now appropriately frame their findings in the Elfn1 knockout mice as indirect evidence, it is worth emphasizing that the study lacks direct in vivo, cell-type-specific recordings and manipulations to more definitively test the proposed mechanism.

      We thank the reviewer for their thorough and constructive evaluation of the manuscript and for highlighting both the conceptual strengths of the study and its technical limitations. We agree that the spatial and cellular specificity of unsorted multi-unit recordings imposes inherent constraints on the interpretation of domain-specific activity, particularly given the narrow width of septal compartments in mice. As now clarified in the manuscript, we do not claim absolute cellular specificity of “septal” recordings but rather interpret them as septal-enriched populations. To directly address this concern, we performed additional threshold-based analysis demonstrating that the key domain-specific effects persist selectively in Layer 4 under stricter spike-detection criteria, supporting a local circuit origin of the critical findings. Further, the more stringent detection criteria (Suppl Fig 3A) collapse the divergence seen in Layer2/3 (Suppl Fig 4C), suggesting that this divergence arises in Layer 4, where SST+ interneuron distributions diverge between barrel and septa.

      We further agree that the Elfn1 knockout results provide indirect, rather than definitive, evidence for causal involvement of SST<sup>+</sup> interneurons and therefore intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, physiological, developmental, and genetic observations. We believe this explicitly moderated interpretation appropriately reflects the scope of the current data while establishing a clear conceptual framework and motivation for future studies employing cell-type-specific recordings and manipulations to directly test the proposed mechanism.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Major comments

      (1) Interpretation of "septal" recordings: The authors claim that the activity recorded from electrodes placed in the septa can be confidently attributed to septal neurons. In my previous review, I raised a major concern that such "septal" recordings likely include spikes from adjacent barrels, given the broad spatial resolution of MUA and the narrowness of the septa in the mouse S1. In fact, the intermediate properties observed in septal recordings from wild-type mice could be explained by a mixture of activity from principal and neighboring barrels-an interpretation that contrasts with the authors' conclusion. Upon reviewing the probe model used (A8x8-Edge-5mm-100-200-177), I noticed a discrepancy between the manufacturer's design and the schematic provided in the manuscript. The electrodes are located near the right edge of the probe rather than the center, suggesting that neurons in adjacent barrels could easily be sampled. In my previous review, I therefore suggested alternative approaches, such as calcium imaging, to more convincingly support the authors' claims. However, the revised manuscript does not include new experiments or additional analyses addressing this issue. Instead, the authors argue that using a high spike detection threshold (SD > 7.5) ensures that recorded activity originates from septal neurons, even though this value does not appear particularly conservative, as it was merely adopted from a previous study without justification in the present context. While I agree that a higher threshold may reduce contamination from distant sources, it does not guarantee that only septal neurons contribute to the signal. By nature, MUA reflects activity from multiple neurons within a radius of at least 50-100 µm. To more rigorously support the claim of spatial specificity, I strongly encourage the authors to reanalyze their existing dataset by systematically varying the spike detection threshold and quantifying how the properties and selectivity of detected units change. If neurons closer to the electrode indeed exhibit distinct domain-specific properties, they should become more prominent as the threshold increases. Such an analysis would strengthen the authors' interpretation and improve the manuscript's impact, even in the absence of new experimental data. Alternatively, the authors could revise their claims to acknowledge that the "septal" electrodes likely record from a population that includes septal neurons as well as neurons located at the periphery of principal and adjacent barrels.

      We agree with the reviewer that, by nature, MUA reflects the activity of multiple neurons within a spatial radius and that recordings obtained from electrodes positioned in the septa may include contributions from neurons located at the periphery of adjacent barrels. This concern is further compounded in superficial layers by probe geometry and orientation: given the narrow width of septa and the lateral spread of processes in upper cortical layers, recordings in L2/3 are inherently more susceptible to spatial mixing than those in layer 4, where columns are more compact and cytoarchitecturally distinct. To directly address these issues, we reanalyzed the same dataset using a more stringent spike detection threshold (SD > 9.5), compared to the originally reported SD > 7.5. Importantly, increasing the threshold selectively reduced or eliminated effects in L2/3, while the key domain-specific differences in L4 responses both the differential MW/SW dynamics in wild-type animals and their attenuation in Elfn1 knockout mice remained robust (the new Supp. Fig. 3. In the manuscript). This threshold-dependent dissociation is consistent with the interpretation that the critical effects reported in L4 arise from neurons spatially closer to the electrode and are less influenced by probe orientation or distant sources, rather than reflecting simple mixing of barrel signals. While this analysis does not claim absolute cellular exclusivity of septal neurons, it provides empirical support that the principal conclusions of the study are robust to stricter spatial sampling criteria and are particularly anchored in L4 circuitry. Accordingly, we now explicitly acknowledge in the manuscript that “septal” recordings likely represent septal-enriched populations rather than purely septal neurons, while emphasizing that the persistence of L4 effects under higher spike-detection thresholds strengthens the conclusion that local L4 inhibitory dynamics underlie the reported functional differences between barrel and septal domains.

      The greater sensitivity of L2/3 results to spike-detection threshold is also expected based on both anatomical considerations and probe geometry. Neurons in L2/3 possess broader horizontal dendritic and axonal arbors and participate in more laterally distributed integration across columns, making population signals in these layers intrinsically less spatially focal. As a result, conservative spike-detection criteria preferentially suppress L2/3 effects, particularly when recordings are obtained with probes optimized for deeper layers. Importantly, our two-photon calcium imaging data while similarly limited to L2/3 demonstrate that SST<sup>+</sup> interneurons show locally measurable and stimulus-specific responses at the single-cell level, providing independent support that L2/3 SST<sup>+</sup> activity is stimulus-modulated rather than artifactual. Taken together, these observations suggest that L2/3 results reflect more distributed and integrative network activity, whereas the L4 effects that persist across thresholds are more directly attributable to local circuit mechanisms. This layer-specific dissociation further supports our interpretation that the central findings of the study are driven by local inhibitory dynamics in L4, with L2/3 activity reflecting downstream integration rather than primary domain-specific computation.

      (2) Interpretation of the Elfn1 KO data: The authors' interpretation that Elfn1-dependent facilitation of SST<sup>+</sup> interneurons underlies the differential sensory responses between barrel and septal domains is conceptually appealing and supported by several converging, albeit indirect, lines of evidence. Specifically, the consistent correspondence among the differential activation of SST<sup>+</sup> neurons upon SWS and MWS, the late development of the barrel-septa differences in the responses to SWS and MWS, and the attenuation of this difference in Elfn1 knockout mice lends plausibility to the proposed model. However, it should be emphasized that the data remain indirect: the study does not include direct recordings of SST<sup>+</sup> neuronal activity from the knockout mice, nor cell-type- specific manipulations to demonstrate causal involvement. The mechanistic explanation therefore represents a hypothesis rather than definitive proof. That said, the authors clearly acknowledge these limitations in the Discussion and appropriately moderate their claims by presenting the SST-Elfn1 mechanism as a working model. Given this careful framing, the current manuscript can be regarded as a valuable conceptual contribution that advances our understanding of how inhibitory dynamics may shape temporal processing in the barrel cortex. Further experiments, as mentioned above, will be essential to test the causal role of this mechanism directly.

      We thank the reviewer for this thoughtful and balanced assessment. We fully agree that the Elfn1 knockout experiments provide indirect rather than definitive evidence for a causal role of SST<sup>+</sup> interneurons in mediating the domain-specific MW/SW response dynamics between barrels and septa For this reason, throughout the revised manuscript we explicitly frame the Elfn1–SST mechanism as a working model rather than a proven mechanism.

      Minor comments:

      The authors have adequately addressed my previous minor comments. In this round, I carefully reviewed the revised manuscript and identified several issues related to references. I would also like to add a brief comment regarding the Discussion section:

      (1) Stachniak et al., 2021 is included in the reference list but is not cited anywhere in the main text. Please either remove this entry or cite it appropriately in the manuscript.

      Removed.

      (2) Yamashita et al., 2018 is cited in the main text (Line 767), but it is not included in the reference list.

      Fixed.

      (3) Sylwestrak and Ghosh, 2012 is cited at Line 261 and Line 270, but likewise absent from the reference list.

      Fixed.

      (4) At Line 497, Chen et al., 2015 is cited, but, the appropriate and original reference would be Chen et al., 2013 (PMID: 23792559), which should either replace or precede the 2015 citation.

      Added.

      (5) At Line 221, El-Boustani et al., 2018 is cited. However, this study is based on the visual cortex, whereas the manuscript concerns the barrel cortex. A more relevant citation (e.g., Lefort et al., 2009 [PMID: 19186171]) would better support the discussion of cellular organization in the barrel cortex. Please consider updating the citation.

      Thank you for this suggestion. We agree with the reviewer and now we have changed El-Boustani with Lefort et al. 2009 as suggested by the reviewer.

      (6) Furthermore, Chakrabarti & Alloway (2006) performed tracer-based mapping of projections from barrel and septal columns in rat S1 and similarly suggested differential organization of M1- and S2projection neurons in the barrel and septal regions.

      Although the current study thoroughly analyzed the layer-specificity of the location of these projection neurons, the lack of explicit discussion of this relevant prior work is a notable omission.

      The authors should incorporate a comparison with these results to better contextualize their findings.

      The following text is added to the discussion: “Our retrograde labeling data supports and expands on previous work proposing similar models (Alloway, 2008; Chakrabarti and Alloway, 2006).”

    1. eLife Assessment

      This potentially valuable study aims to investigate neural correlates of spatial attention in whisker somatosensory cortex (S1) in mice, finding increased sensory-evoked spiking when the mice appear to be attending to the contralateral whiskers. Although some of the results appear to be robust despite relatively small effect sizes, overall the findings are incompletely supported, because attentional modulation is insufficiently distinguished from learning of stimulus-response contingencies, and because the analyses do not adequately consider orofacial movements that may contribute key confounds.

    2. Reviewer #1 (Public review):

      The paper uses a passive whisker detection task in mice to identify a behavioral phenomenon that can reasonably be interpreted as spatial attentional capture. The attentional effect occurs transiently after a successful whisker stimulus detection yields reward, and lasts for a few trials before subsiding. The attentional effect is to the right or left whiskers, depending on whether right or left whiskers are rewarded; no finer spatial resolution for attention was tested. By recording whisker-evoked spiking from single units in S1, the authors show that this form of spatial attention increases the gain of whisker-evoked neuronal responses in S1 for a large subset of S1 units. In contrast, neural responses are not modulated by overall task engagement. Together, these findings show a neural signature of spatial attention in S1 cortex. Because whisker or facial movements were not tracked, it is not clear whether this represents covert attention or whisker movement in response to previously rewarded stimuli, which would be a form of overt attention.

      Substantial attentional modulation of neural responses was observed for a subset of whisker-responsive S1 units, but the effect size was small on average for the total unit population. The top 25% of units showed a ~12% attentional response modulation (relative to firing rate range for each unit), but the median unit showed only a 1.3% response modulation. It would have been useful to analyze the magnitude or prevalence of attentional modulation across layers or in fast-spiking vs. regular spiking units, but this was not reported.

      Major

      (1) It is hard to interpret the underlying causes of the attentional modulation of neural activity without having measured whisker and facial movement. This is a particular issue in S1, where whisker movement against the stimulation grid can alter the mechanical efficiency of stimulus delivery. Such movements would represent overt attention, which would engage an entirely different neural mechanism than covert attention.

      (2) An interesting debate is whether the behavioral phenomenon is best described as attention or as dynamic learning of the stimulus-response association for that block. In Posner-type cued attention tasks, and also in many block-type attention tasks in rodents, animals receive reward for successfully detecting either cued or uncued stimuli, and thus attention (higher response probability or improved psychometric sensitivity for cued stimuli) is at least partially dissociated from the stimulus-reward contingency. That is not the case here. The fact that mice have difficulty learning the contingency reversal suggests that the phenomenon is better explained by attention than by learning the contingency; however, to prove this clearly, the existence of the attentional effect on neural activity in Block 1 vs. Block 2 would have to be shown.

      (3) Some of the graphical representations of the attentional modulation of neural activity are unclear. The single-unit example of attentional modulation is quite strong (Figure 3d). The mean response for the top 25% of units is also visually clear (Figure 3f). But the effect is not apparent at all in Figure 3e, which the figure legend says shows every unit. What is the yellow point and line in this figure? Why isn't the attentional effect visible in this panel? Perhaps I am misunderstanding Figure 3e, but it is not clear to me why it compares Pref>0.5 to Pref<0.5, when the intended analysis suggests it should be Pref>0 to Pref<0? Also in Figure 3, it is critical for the reader to know whether panels 3g-3h represent the top 25% of units or all units. Neither the results text nor the legend is clear on this.

      (4) There is a missed opportunity to quantify attentional modulation across cortical layers, since laminar probes and Neuropixels probes were used for the recordings. In addition, there is no separation of fast-spiking from regular-spiking units, and no quantitative metrics are provided to assess the quality of single units. This could reveal key aspects of cortical processing of attentional signals.

    3. Reviewer #2 (Public review):

      Summary:

      Dyce et al investigate the modulation of sensory responses in the somatosensory 'barrel' cortex during a novel whisker vibration detection task in head-fixed mice, aiming to find correlates of spatial attention in both the animals' behavior and their neuronal activity.

      Strengths:

      The authors produced an extensive and parameterized dataset of both behavioral responses and neuronal activity, with >3000 single units of which >1400 were responsive.

      Weaknesses:

      In my view, the main conclusions of the manuscript are not currently well supported by the data.

      The authors effectively define "spatial attention" as a state where an animal responds more to a stimulus that gives more rewards (out of two possible stimuli presented on different sides of the snout, i.e., segregated spatially). If one defines spatial attention purely in these terms, then their findings do show neuronal correlates of spatial attention. However, those neuronal correlates can be explained by known aspects of neuronal responses in the barrel cortex.

      This plays out in several different ways:

      From the behavioral point of view, greater attention may correlate with an increased hit rate to stimuli on the rewarded side, but in the absence of other supporting measurements, the relationship could well be the opposite: an animal could pay more (rather than less) attention to the stimulus delivered on the unrewarded side, to make sure it suppresses the incorrect response. It is impossible to tell, as the data don't provide an independent measurement of whether the animal is paying greater attention to, or is more aware of, one side than the other, nor do they provide an independent measurement of neuronal tuning on either side. There is no separate measurement of arousal either (e.g., via pupillometry or locomotion).

      The experimental design involved two blocks on each daily task session, with the second block reversing the side on which rewarded stimuli were delivered. Reinforcing one's doubts about the behavior and its interpretation, mice had much poorer performance on each day's second block, to the extent that perceptual sensitivity (d') was the same for both sides: d' did not increase after reward reversal for stimuli on the initially unrewarded side. This further emphasizes the lack of a separate demonstration of focused "spatial attention".

      Much of the data (both behavioral and neuronal) could be accounted for, e.g., by a strategy where the mouse keeps a token in working memory of what side seems to be driving rewards, while maintaining equally strong sensory drive on both sides, but with no attentional shift at all. The policy would be to respond more whenever the stimulated side matches the token in memory (thus also reinforcing the token, thus enhancing performance next time). This would be easily implemented with a disinhibitory reward-modulation signal such as the one multiple researchers have found carried by VIP neurons (e.g., Szadai et al DOI: 10.7554/eLife.78815).

      Similarly, the fact that "attended trials" (Pref > 0) produced greater responses than "unattended trials" appears to be explainable as follows. Here, "attended" trials are those where the contralateral stimulus is presented (and, if responded to, is rewarded), "unattended" trials are those where the stimulus is ipsilateral (and not rewarded). The animal responds more (at least in the first block) to stimuli delivered to the contralateral pad - i.e., rewarded as opposed to unrewarded ones. Beyond the knowledge mentioned above that cortex-wide VIP sensitivity to rewards can drive disinhibition in general, activity modulation dependent on rewards and outcomes (and stimulus value) has been established specifically in the barrel cortex (e.g., Lacefield et al DOI: 10.1016/j.celrep.2019.01.093, Bale et al DOI: 10.1016/j.cub.2020.10.059, Banerjee et al DOI: 10.1038/s41586-020-2704-z, Chereau et al 10.1038/s41467-020-17005-x). The reward- and value-evoked activity demonstrated in those papers would suffice to predict more activity at the contralateral electrode on "attended" trials, along the lines of the findings in Ramamurthy et al (DOI: 10.1038/s41467-025-60592-w) and consistent also with the enhanced "attentional modulation" on hit trials.

      Other aspects of the analysis and terminology lead to confusing outcomes. For example, in the analysis in Figure 3, Performance averaged in a set of trials around a given trial is defined as the mean rate of responses to stimulation on either side - regardless of whether those responses are correct (since the stimuli can be on either side, but only one side is correct and gets rewarded and putatively reinforced). Thus, this definition of "Performance" can increase with the rate of incorrect licks to the wrong side and is at odds with the normal use of the word. On trials where this Perf = 1 and the stimuli are balanced on either side, this corresponds to a true performance (and reward rate) of only 0.5 - what one would normally consider random discrimination between the sides. Thus, Perf = 1 trials may still give a low reward rate and, if responses scale with reward, a small effect of reward. Hence, based on known properties of reward dependence, greater correlation of neuronal activity with "Preference" than with "Performance" would be expected, rather than reflecting a new aspect of "spatial attention". A definition of performance more in line with established practice and measuring side-to-side discrimination (corresponding more closely to the authors' "Preference" parameter) would have shown this more clearly.

    4. Author response:

      (1) Introduction & Roadmap

      We are grateful to the Reviewers for engaging with outstanding questions relating to our findings’ connections to multiple subdisciplines of cognitive neuroscience. Noting that Reviewers 1 and 2 interpreted our findings differently, we welcome the opportunity to engage in what Reviewer 1 characterised as “an interesting debate”. To promote a shared understanding and discussion of our findings, we have organised our response to address more technical comments first.

      Our provisional response is organised as follows: Section 2 addresses selected technical comments relating to our Results. Section 3 addresses comments related to the design of our behavioural paradigm. Section 4 focuses on the broader interpretation of our findings. Section 5 concludes our provisional response with potential future directions and a summary of the significance of our findings.

      (2) Selected technical comments related to our Results

      We apologise to Reviewer 2 for the confusion in relation to the meaning of “attended” and “unattended” trials. What we said was “Positive Pref values indicate a higher response rate to the contralateral side than the ipsilateral side (relative to the electrode)” (Figure 3c caption), “we indexed all contralateral whisker vibrations according to their associated Perf and Pref” (Results text), and “we divided trials into (contralaterally) attended (Pref<sub>C/L</sub>: Pref>0) and unattended (Pref<sub>I/L</sub>: Pref<0) groups” (Results text). We can confirm that we defined an “unattended trial” (Pref<0) as a contralateral stimulus trial in the centre of an epoch (10-15 trials) within which the mouse responded (licked) more frequently to ipsilateral stimuli. Critically, we did not define an unattended trial as an ipsilateral stimulus trial. Furthermore, attention thus defined (i.e. Pref>0) can vary independently of the whisker stimulus associated with rewards. Indeed, while we initially did not include this result in our paper for the sake of brevity, even unrewarded “attended” trials (Pref>0) evoked significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials. We note that this is an analysis suggested by Reviewer 1, and we will include and discuss this result in our revised manuscript (e.g. in relation to literature suggested by Reviewer 2). For additional clarity, we use “Performance” (Perf) in relation to overall stimulus detection, consistent with the analysis of Lee et al. (2020), which found this measure was correlated with pupil diameter in a vibrissal target detection task.

      We thank Reviewer 1 for noticing that the axes on Figure 3e should be labelled “Pref>0” (Y axis) and “Pref<0” (X axis), as suggested by the figure caption. We will correct this in our revised submission. The yellow point on Fig 3e shows the unit from Fig 3d, while the yellow line in Fig 3e shows the magnitude of that unit’s (non-normalised) gain modulation. While this is alluded to in the Results text (“The example unit in Figure 3d is in the 93rd percentile of units for raw modulation depth (ΔHits(attended – unattended) = 3.3 spikes/second; yellow line in Fig.3e)”, this should be explained in the Figure caption, and it will be in our revised manuscript. We would also like to clarify that Figures 3g–3h display results for all units, not just the top 25%. We agree this is not sufficiently clear and we will rectify this in our revised manuscript. Addressing Reviewer 2, while we acknowledge that mice responded less to both stimuli in the second block, they also meaningfully adjusted their behaviour to the reversal in reward contingencies: their responses to the previously rewarded stimulus reduced significantly more than those to the previously unrewarded stimulus.

      (3) Design of the behavioural paradigm

      We made a deliberate design choice to maximise the ecological validity of our behavioural paradigm, and note that there are advantages to doing so. For example, our paradigm can be used to show that even unrewarded “attended” trials (Pref>0) evoke significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials (see Section 2, above). Indeed, it is precisely this finding that makes our paradigm uniquely suited to the investigation of value-driven attentional capture (Anderson et al., 2011): in this instance attention directed to stimuli that are no longer rewarded despite equal availability of rewarded stimuli. This finding also demonstrates that our paradigm dissociates attention from stimulus-reward contingency at least as well as other paradigms which have been successfully used to study spatial attention in mice. As noted in Section 2, we will discuss this result in relation to other relevant research (e.g. Ramamurthy et al., 2025) in our revised manuscript.

      Briefly, the direct manipulation of reward contingencies is one of two noteworthy methodological distinctions between our own paradigm and that of Ramamurthy and colleagues (2025). The task of Ramamurthy et al. (2025) associated all whisker stimuli with rewards and delivered stimuli to different whiskers on a single whisker pad. These methodological distinctions may have reduced the relevance of the spatial differences between stimuli to the mice undertaking the task. Indeed, it is not certain that a mouse would treat the unilateral variation in whisker stimulation Ramamurthy and colleagues delivered as primarily spatial or featural. The psychophysical and neural differences between spatial and featural attention in humans suggest dissociable underlying mechanisms, and the same may be true in mice. Thus, our own paradigm may more effectively isolate spatial attention from featural attention. Conversely, to the extent that the findings of Ramamurthy and colleagues do reflect spatial attention, our combined findings and paradigms help elucidate the associated mechanisms across spatial scales in mice.

      We acknowledge that spatial cueing is well-suited to isolating the effects of covert attention from other forms of attention. However, it should be noted that spatial cueing in rodents is subject to its own challenges, including limitations in trial numbers due to the required manipulation of stimulus intensity (Reynolds et al., 2000; Herrmann et al., 2010), cue validity and associated trial probabilities (Peterson & Gibson, 2011; Girardi et al., 2013). Such experiments are further complicated by the duration and efficacy of training (i.e. the number of mice that learn the task; Wang & Krauzlis, 2018; Hu & Dan, 2022). It is also worth noting that trial probability manipulations introduce the same limitation in trial numbers with block-type attention tasks (You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022).

      While there are clear differences between our own paradigm and those mentioned above, there are also important similarities. First, these tasks are all goal-directed, stimulus-driven, and reliant on learned task contingencies (e.g. Peterson & Gibson, 2011; Girardi et al., 2013). Furthermore, these paradigms are all operant conditioning protocols which leverage learned stimulus-reward contingencies to train attention-related behaviours in mice. A noteworthy similarity between our findings and those of authors using block-type attention tasks in particular (e.g. You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022) is the observation of apparent attentional biases in behavioural responses independent of the experimental manipulations (i.e. stimulus probability / reward contingency).

      (4) Comments relating to the broader interpretation and discussion of our findings

      Fundamentally, attention involves dedicating limited processing resources to some stimulus events at the expense of others. The design of our behavioural paradigm was informed by existing literature on spatial attention in humans, non-human primates, and mice. Our choice of behavioural and neuronal measures as proxies for attention in mice is consistent with this literature. It is technically possible “an animal could pay ‘more’ (rather than less) attention to the stimulus delivered on the unrewarded side, to make sure it suppresses the incorrect response”, but this seems unlikely given what is known about how attention is typically allocated in such tasks, based on the previously mentioned literature.

      With respect to the interpretation and discussion of our findings, Reviewer 1 describes them as “a behavioral phenomenon that can reasonably be interpreted as spatial attentional capture” but suggests they do not clearly distinguish whether this attentional capture is covert or overt. We respectfully disagree for three reasons. First, as discussed in our paper, whisker motion during detection tasks has consistently been associated with reduced detection performance (Ollerenshaw et al., 2012; Kyriakatos et al., 2017; Vandevelde et al., 2023), suggesting that a “receptive” strategy (Diamond & Arabzadeh, 2013) of whisker immobilisation is more applicable to the current data than a “generative” strategy of asymmetric whisker movement (O'Connor et al., 2010; Dominiak et al., 2019). Second, if our behavioural and neuronal findings were due to the mice moving their whiskers to maximise contact with the meshes, we would expect increased evoked neuronal responses to be associated with greater Perf, not just with greater Pref. This pattern was not observed. Of course, the mice might have employed different whisker movement strategies during epochs of high Pref and Perf, but this seems unlikely and is not a parsimonious explanation for our findings. Third, as noted in the Methods section of the paper, we deliberately positioned the meshes close to the base of the whiskers, limiting the impact of whisker movements on stimulus detectability and the incentive to make them.

      In contrast, Reviewer 2 questions the interpretation of our findings as evidence of spatial attention and suggests they might reflect working memory instead. Current research suggests attention and working memory are intimately related integrative brain functions. Indeed, some researchers have even proposed that working memory might be a form of internally directed attention (Awh & Jonides, 2001; Chun, 2011; Gazzaley & Nobre, 2012; Kiyonaga & Egner, 2013; or vice versa: Libedinsky & Fernandez, 2019). Consistent with the comments of Reviewer 2, more recent work seems to emphasise the coordination of attention and working memory (e.g. Joe & Kim, 2023; Zhu et al., 2026; for reviews see Huynh Cong & Kerzel, 2021; van Ede & Nobre, 2023), along with shared mechanisms (Kiyonaga et al., 2021; Panichello & Buschman, 2021), and nuanced dissociations (Liu et al., 2025). Attention is difficult to dissociate from working memory partly because there are multiple definitions (and/or types) of attention. We did not discuss the various definitions and/or forms of attention at length in our paper, but we will briefly discuss this in the revised manuscript.

      The “interesting debate” to which Reviewer 1 refers could also be described as vigorous, despite approximately three decades of research. This debate broadly relates to the degree to which attentional control is driven by exogenous (e.g. colour contrast) versus endogenous factors (e.g. the focus of spatial attention, see Fig.2 in Belopolsky et al., 2007; see also: Liesefeld & Mueller, 2020; Manini et al., 2021; Beffara et al., 2022), and the degree to which this is a function of experimental context. The review article by Luck et al. (2021) entitled “Progress toward resolving the attentional capture debate” provides a striking illustration of this debate, as do the twenty-two commentaries (and three commentary responses) associated with it. Admittedly, this debate largely revolves around human attention experiments, and human cognition may be more complex than mouse cognition. However, the complexity of human cognition may also be easier to study and appreciate because complex behavioural experiments can be explained to, understood, and performed by human participants with relative ease.

      (5) Comments relating to future directions and the significance of our findings

      The complexity of the attentional capture debate underscores the importance of developing accessible and scalable animal experiments which can be used to provide mechanistic insights. If the human attention literature is any indication, a diversity of rodent experimental paradigms will be necessary to thoroughly map the neuronal implementation of spatial attention. Returning to our paradigm, Reviewer 1 noted that valuable insights into the mechanisms of vibrissal spatial attention might be obtained from comparing the magnitude of attentional modulation we observed between putative regular and fast-spiking categories of units, and between units located in different cortical layers. We agree it is important to understand spatial attention with cell-type and circuit (including laminar) specificity. However, because we could not persuasively cluster our units based on waveform width, and because of the lack of histological data, segregating units on the basis of such variables is not feasible. Despite our assertion that our findings reflect the effects of covert attention (contra Reviewer 1), we agree that future experiments will be required to conclusively rule out overt attention. Noting the proximity of the meshes to the base of the whiskers in our paradigm, and the difficulty of tracking whiskers in this context, Botulinum toxin injections (as in Ramamurthy et al., 2025) might be a means of achieving this.

      The above notwithstanding, our findings provide multiple contributions to the literature on spatial attention (and perhaps working memory). We detected significant attentional gain modulation across a population of 1461 responsive units. While the gain modulation exhibited by the median unit was modest (albeit statistically significant), the top 25% of responsive units showed a ~12% response modulation (relative to firing rate range for each unit), and ~21% of responsive units were suppressed by the average vibrissal stimulus in the unattended state. Our experimental framework offers an accessible platform for future studies leveraging genetic and circuit-level interventions to dissect the cell-type specific mechanisms of spatial attention. Our work is timely, noting the recent focus of human research on the nexus of attention, selection history, and valence (e.g. Serences, 2008; Della Libera & Chelazzi, 2009; Della Libera et al., 2011; van den Berg et al., 2014; Kim & Anderson, 2019, 2023). Our work is also uniquely poised to stimulate new interdisciplinary research into the circuit mechanisms of value-driven attentional capture, with translational relevance to psychopathologies such as ADHD, addiction, and depression; where value-driven attentional capture is altered (for a review see Anderson, 2021).

      References

      Anderson, B. A. (2021). Relating value-driven attention to psychopathology. Curr Opin Psychol, 39, 48-54. https://doi.org/10.1016/j.copsyc.2020.07.010

      Anderson, B. A., Laurent, P. A., & Yantis, S. (2011). Value-driven attentional capture. Proceedings of the National Academy of Sciences of the United States of America, 108(25), 10367-10371. https://doi.org/10.1073/pnas.1104047108

      Awh, E., & Jonides, J. (2001). Overlapping mechanisms of attention and spatial working memory. Trends Cogn Sci, 5(3), 119-126. https://doi.org/10.1016/s1364-6613(00)01593-x

      Beffara, B., Hadj-Bouziane, F., Ben Hamed, S., Boehler, C. N., Chelazzi, L., Santandrea, E., & Macaluso, E. (2022). Dynamic causal interactions between occipital and parietal cortex explain how endogenous spatial attention and stimulus-driven salience jointly shape the distribution of processing priorities in 2D visual space. Neuroimage, 255. https://doi.org/10.1016/j.neuroimage.2022.119206

      Belopolsky, A. V., Zwaan, L., Theeuwes, J., & Kramer, A. F. (2007). The size of an attentional window modulates attentional capture by color singletons. Psychonomic Bulletin & Review, 14(5), 934-938. https://doi.org/10.3758/Bf03194124

      Chun, M. M. (2011). Visual working memory as visual attention sustained internally over time. Neuropsychologia, 49(6), 1407-1409. https://doi.org/10.1016/j.neuropsychologia.2011.01.029

      Della Libera, C., & Chelazzi, L. (2009). Learning to Attend and to Ignore Is a Matter of Gains and Losses. Psychological Science, 20(6), 778-784. https://doi.org/10.1111/j.1467-9280.2009.02360.x

      Della Libera, C., Perlato, A., & Chelazzi, L. (2011). Dissociable Effects of Reward on Attentional Learning: From Passive Associations to Active Monitoring. PLoS One, 6(4). https://doi.org/10.1371/journal.pone.0019460

      Diamond, M. E., & Arabzadeh, E. (2013). Whisker sensory system - from receptor to decision. Prog Neurobiol, 103, 28-40. https://doi.org/10.1016/j.pneurobio.2012.05.013

      Dominiak, S. E., Nashaat, M. A., Sehara, K., Oraby, H., Larkum, M. E., & Sachdev, R. N. S. (2019). Whisking Asymmetry Signals Motor Preparation and the Behavioral State of Mice. J Neurosci, 39(49), 9818-9830. https://doi.org/10.1523/JNEUROSCI.1809-19.2019

      Gazzaley, A., & Nobre, A. C. (2012). Top-down modulation: bridging selective attention and working memory. Trends Cogn Sci, 16(2), 129-135. https://doi.org/10.1016/j.tics.2011.11.014

      Girardi, G., Antonucci, G., & Nico, D. (2013). Cueing spatial attention through timing and probability. Cortex, 49(1), 211-221. https://doi.org/10.1016/j.cortex.2011.08.010

      Herrmann, K., Montaser-Kouhsari, L., Carrasco, M., & Heeger, D. J. (2010). When size matters: attention affects performance by contrast or response gain. Nat Neurosci, 13(12), 1554-1559. https://doi.org/10.1038/nn.2669

      Hu, F., & Dan, Y. (2022). An inferior-superior colliculus circuit controls auditory cue-directed visual spatial attention. Neuron, 110(1), 109-119 e103. https://doi.org/10.1016/j.neuron.2021.10.004

      Huynh Cong, S., & Kerzel, D. (2021). Allocation of resources in working memory: Theoretical and empirical implications for visual search. Psychon Bull Rev, 28(4), 1093-1111. https://doi.org/10.3758/s13423-021-01881-5

      Joe, J., & Kim, M. S. (2023). Spatial Attention in Visual Working Memory Strengthens Feature-Location Binding. Vision (Basel), 7(4). https://doi.org/10.3390/vision7040079

      Kanamori, T., & Mrsic-Flogel, T. D. (2022). Independent response modulation of visual cortical neurons by attentional and behavioral states. Neuron, 110(23), 3907-3918 e3906. https://doi.org/10.1016/j.neuron.2022.08.028

      Kim, H., & Anderson, B. A. (2019). Dissociable neural mechanisms underlie value-driven and selection-driven attentional capture. Brain Research, 1708, 109-115. https://doi.org/10.1016/j.brainres.2018.11.026

      Kim, H., & Anderson, B. A. (2023). Primary Rewards and Aversive Outcomes Have Comparable Effects on Attentional Bias. Behavioral Neuroscience, 137(2), 89-94. https://doi.org/10.1037/bne0000543

      Kiyonaga, A., & Egner, T. (2013). Working memory as internal attention: toward an integrative account of internal and external selection processes. Psychon Bull Rev, 20(2), 228-242. https://doi.org/10.3758/s13423-012-0359-y

      Kiyonaga, A., Powers, J. P., Chiu, Y. C., & Egner, T. (2021). Hemisphere-specific Parietal Contributions to the Interplay between Working Memory and Attention. J Cogn Neurosci, 33(8), 1428-1441. https://doi.org/10.1162/jocn_a_01740

      Kyriakatos, A., Sadashivaiah, V., Zhang, Y., Motta, A., Auffret, M., & Petersen, C. C. (2017). Voltage-sensitive dye imaging of mouse neocortex during a whisker detection task. Neurophotonics, 4(3), 031204. https://doi.org/10.1117/1.NPh.4.3.031204

      Lee, C. C. Y., Kheradpezhouh, E., Diamond, M. E., & Arabzadeh, E. (2020). State-Dependent Changes in Perception and Coding in the Mouse Somatosensory Cortex. Cell Rep, 32(13), 108197. https://doi.org/10.1016/j.celrep.2020.108197

      Libedinsky, C. D., & Fernandez, P. F. (2019). Graded Memory: A Cognitive Category to Replace Spatial Sustained Attention and Working Memory
 Yale J Biol Med, 92(1), 121-125. https://www.ncbi.nlm.nih.gov/pubmed/30923479

      Liesefeld, H. R., & Mueller, H. J. (2020). A theoretical attempt to revive the serial/parallel-search dichotomy. Attention Perception & Psychophysics, 82(1), 228-245. https://doi.org/10.3758/s13414-019-01819-z

      Liu, Y., Fu, Y., Tang, E., Wu, H., Han, J., Xie, M., Zhang, Y., Peng, B., Huang, J., Liu, H., Chen, H., & Qin, P. (2025). Neural dissociation of attention and working memory through inhibitory control. Nat Commun, 17(1), 22. https://doi.org/10.1038/s41467-025-66553-7

      Luck, S. J., Gaspelin, N., Folk, C. L., Remington, R. W., & Theeuwes, J. (2021). Progress toward resolving the attentional capture debate. Visual Cognition, 29(1), 1-21. https://doi.org/10.1080/13506285.2020.1848949

      Manini, G., Botta, F., Martin-Arevalo, E., Ferrari, V., & Lupianez, J. (2021). Attentional Capture From Inside vs. Outside the Attentional Focus. Frontiers in Psychology, 12. https://doi.org/10.3389/fpsyg.2021.758747

      O'Connor, D. H., Clack, N. G., Huber, D., Komiyama, T., Myers, E. W., & Svoboda, K. (2010). Vibrissa-based object localization in head-fixed mice. J Neurosci, 30(5), 1947-1967. https://doi.org/10.1523/JNEUROSCI.3762-09.2010

      Ollerenshaw, D. R., Bari, B. A., Millard, D. C., Orr, L. E., Wang, Q., & Stanley, G. B. (2012). Detection of tactile inputs in the rat vibrissa pathway. J Neurophysiol, 108(2), 479-490. https://doi.org/10.1152/jn.00004.2012

      Panichello, M. F., & Buschman, T. J. (2021). Shared mechanisms underlie the control of working memory and attention. Nature, 592(7855), 601-605. https://doi.org/10.1038/s41586-021-03390-w

      Peterson, S. A., & Gibson, T. N. (2011). Implicit attentional orienting in a target detection task with central cues. Conscious Cogn, 20(4), 1532-1547. https://doi.org/10.1016/j.concog.2011.07.004

      Ramamurthy, D. L., Rodriguez, L., Cen, C., Li, S., Chen, A., & Feldman, D. E. (2025). Reward history guides focal attention in whisker somatosensory cortex. Nat Commun, 16(1), 5580. https://doi.org/10.1038/s41467-025-60592-w

      Reynolds, J. H., Pasternak, T., & Desimone, R. (2000). Attention increases sensitivity of V4 neurons. Neuron, 26(3), 703-714. https://doi.org/10.1016/s0896-6273(00)81206-4

      Serences, J. T. (2008). Value-Based Modulations in Human Visual Cortex. Neuron, 60(6), 1169-1181. https://doi.org/10.1016/j.neuron.2008.10.051

      van den Berg, B., Krebs, R. M., Lorist, M. M., & Woldorff, M. G. (2014). Utilization of reward-prospect enhances preparatory attention and reduces stimulus conflict. Cognitive Affective & Behavioral Neuroscience, 14(2), 561-577. https://doi.org/10.3758/s13415-014-0281-z

      van Ede, F., & Nobre, A. C. (2023). Turning Attention Inside Out: How Working Memory Serves Behavior. Annu Rev Psychol, 74, 137-165. https://doi.org/10.1146/annurev-psych-021422-041757

      Vandevelde, J. R., Yang, J. W., Albrecht, S., Lam, H., Kaufmann, P., Luhmann, H. J., & Stuttgen, M. C. (2023). Layer- and cell-type-specific differences in neural activity in mouse barrel cortex during a whisker detection task. Cereb Cortex, 33(4), 1361-1382. https://doi.org/10.1093/cercor/bhac141

      Wang, L., & Krauzlis, R. J. (2018). Visual Selective Attention in Mice. Curr Biol, 28(5), 676-685 e674. https://doi.org/10.1016/j.cub.2018.01.038

      You, W. K., & Mysore, S. P. (2020). Endogenous and exogenous control of visuospatial selective attention in freely behaving mice. Nat Commun, 11(1), 1986. https://doi.org/10.1038/s41467-020-15909-2

      Zhu, P., Guan, C., Fu, Y., Shen, M., & Chen, H. (2026). Working memory encoding of attended information is adaptive to future relevance. J Exp Psychol Learn Mem Cogn. https://doi.org/10.1037/xlm0001582

    1. eLife Assessment

      This valuable study compares hippocampal-cortical functional connectivity to various other brain measures and examines their development across youth. It uses sophisticated analyses replicated in multiple datasets, but provides incomplete evidence to support the primary claim that hippocampal-cortical connectivity relates to cognitive maturation. The manuscript would benefit from a more nuanced consideration of the biological basis of some of the derived imaging measures and the limitations of the cross-sectional design. This work will be of interest to neuroimaging specialists and cognitive neuroscientists.

    2. Reviewer #1 (Public review):

      Summary:

      The authors studied the development of hippocampal connectivity gradients based on open datasets and performed correlation analyses with other MRI features as well as gene expression information from other datasets. Although the main findings are correlational and cross-sectional, the analyses are overall sophisticated and replicated in several datasets.

      Strengths:

      The hippocampus is a key region in understanding large-scale brain organization and cognition, and the authors applied advanced and suitable analytics to study its development. The paper is overall well-organized and well-written, and the findings are relevant for studying large-scale brain development.

      Weaknesses:

      While sophisticated, several of the analyses appear mainly correlational, cross-sectional, and rely on cross-dataset contextualization, which should also be stated as a limitation of the current work.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors aim to assess how the functional organisation of the hippocampus is related to the geometry and neurobiological differences of the hippocampus. In particular, the authors focus on the first three eigenvectors of hippocampal-cortical functional connectivity, based on non-linear dimensionality reduction on resting-state functional MRI data. Furthermore, the work aims to describe changes in these functional axes and their relation to other factors throughout youth and evaluate whether they are predictive of individual variations in cognition.

      Strengths:

      A major strength of this study is the attempt to replicate key findings across multiple developmental cohorts.

      Weaknesses:

      The major weaknesses of the manuscript center on gaps in technical transparency and several conceptual inaccuracies. The machine learning methodology used for cognitive prediction is scarce, leaving little means to evaluate whether the behavioral results suffer from data leakage or overfitting. The introduction sets up an oversimplified historical premise regarding the field's understanding and appreciation of hippocampal connectivity, and contains several incorrect references that throw doubt on the argumentation. Additionally, T1w/T2w signal intensity is incorrectly used as synonymous with myelin, despite gold-standard histological validation showing a non-significant correlation between T1w/T2w and myelin staining (Sandrone et al., 2013).

      Appraisal of Aims and Conclusions:

      The authors partially achieve their aims by illustrating certain age-related changes in hippocampal function; however, the correlative study design is not equipped to examine how these changes are "shaped" by geometry, myelination, or gene expression (especially the latter two). Furthermore, conclusions were often overstated based on small effect sizes.

      Context and Field Impact:

      This work adds to a growing body of literature focused on gradient-based representations of hippocampal topology. By applying these methods across a wide developmental age bracket, it provides a useful reference point for how the hippocampus and wider cortex interact during maturation. To improve utility to the neuroimaging and cognitive neuroscience communities, the nesting of subfields within the eigenvector topology should be addressed, too.

    1. eLife Assessment

      This valuable manuscript investigates how Drosophila larvae make foraging decisions in patchy environments with controlled resource density and valence; using movement tracking in bounded arenas, the authors show that larvae's patch residence time (PRT) differs depending on resource type, environmental context, and prior experience. A drift-diffusion model is used to describe patch-leaving behaviour, suggesting that an integration process may underlie stay-leave decisions during foraging. The strength of the evidence is mostly solid, but the interpretation and use of PRT needs further investigation, as PRT could be a direct effect of resource concentration on locomotion. Explicit reports of PRT statistical tests are needed for rigorous interpretation.

    2. Reviewer #1 (Public review):

      Summary:

      Mudunuri et al. investigate the foraging response of Drosophila larvae in response to patchy resources of distinct value (concentration of nutrient or valence). They show that larvae adjust their behavior according to both the quality and valence of available resources. Interestingly, previous exposure to resources of lower value increases the permanence time in resources of greater value. This suggests that larvae can value, remember and adapt their behaviour in response to previous foraging experience.

      They perform a simple integration model that recapitulates the larval behaviour.

      Strengths:

      This paper uses a very well-controlled foraging set-up where larvae are tested individually and for 3 hours, allowing for a good statistical analysis of their behaviour.

      They investigate for the first time the ability of Drosophila larvae to perceive, remember and compare the quality and valence of distinct resources. It is very exciting, as it will open up the field of foraging decision studies using the fruitfly larvae.

      Weaknesses:

      (1) Most of the analysis depends on the thresholding, but it is not clear what increasing the radius of analysis means in terms of foraging. There are two issues here:

      a) What is the behaviour of the larvae on the edges of the patch? It is obvious that the fructose or the NaCl will diffuse at the edge, so are they remaining in the proximity because they are actively feeding (exploiting) on this decaying concentration, or are they sensing the lower gradient and they are actually looking (chemosensing) for the higher concentration? The behaviour at the edge is really different (check sucrose in Wosniack et al. 2022), and there might be a way of avoiding the diffusion by actually adding a plastic ring and pouring the agar + resource in there. The effect of the ring, per se, would still have to be tested.

      b) How was the threshold selected? It is very likely that the concentration at the patch boundary will be very different for 1M and 0.1 M. Could the authors explain why they chose such a distance? What does majority of larvae mean? Is the "majority" the same for 0.1M and 1M? Is there a relationship between the threshold chosen and the diffusion of fructose and NaCl?

      (2) The word exploitation is used in the paper, but there are many instances where it is unclear whether that is the case. This should be clarified since there are no controls for exploitation.

      (3) In the experiments analysing the adaptation of foraging behaviour, it is not clear if the first and second patch means that only 2 patches were analysed per larva or the first and second in a sequence of patches visited. I think it is the second option (because of Figure S3D), but the authors should clarify this. Also, we do not know how many animals were tested. The number of data points in 4C (4G) compared to 4D (4H) seems very different.<br /> Regarding the results, which are very interesting, why aren't the larvae spending less time in the 0.1M sucrose patch after having fed on a 1M patch, while they spend more time in a 1M after a 0.1M? Could it be that the difference in residence time is correlated with their hunger rather than the comparison between conditions?

      (4) I am not an expert in this type of model, and I would appreciate it if the authors could explain how the values of the drift and leak have been fitted in Figure 5H. If possible, I would recommend adding a graph showing the parameter exploration of distinct possible combinations of values.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how Drosophila larvae make foraging decisions in patchy environments with controlled resource density and valence. Using movement tracking in bounded arenas, the authors show that larvae's patch residence time (PRT) differs depending on resource type, environmental context, and prior experience.

      The authors vary whether the environment is homogenous (all patches are equal) or heterogenous (mixed patches) and whether a higher density of the resource is appetitive (food) or aversive (salt). The most salient results are that in heterogeneous environments, larvae remain longer on higher-density patches of fructose, while they stay shorter in higher-density salt patches. The study further demonstrates that prior foraging experience influences subsequent patch residence time (PRT).

      A drift-diffusion model is used to describe patch-leaving behavior, suggesting that an integration process may underlie stay-leave decisions during foraging. Overall, the work provides a useful behavioral system for studying foraging behaviour and highlights the role of context and experience in shaping larval foraging strategies.

      Strengths:

      A major strength of the manuscript is the behavioral system. The assay is simple, well-controlled, and suitable for realistic spatial and temporal scale tracking of individual larvae. The use of non-volatile resources and embedded patches minimizes confounds from olfactory navigation and allows the authors to focus on local patch exploitation, return behavior, and experience-dependent decisions.

      The results regarding patch resident time (how long larvae stay in patches of different resource density) are convincing. In homogeneous environments, larvae spend more time on patches with a higher density of food (0.1M > 0.01M) and less time in patches with a lower density of salt (0.01M > 0.1M), indicating that their behaviour is sensitive to the valence of the resource. Further, larvae do not simply respond to current circumstances, since PRT in a given patch is sensitive to the quality of the preceding one encountered, showing some kind of memory.

      Weaknesses:

      (1) The theoretical background of the experiment, as exposed in the Introduction, is somewhat misleading. The experiment is based on patches of sufficient size for the individual larvae not to deplete them through their activity, so that the intake rate is constant while exploiting a given patch. In those circumstances, the theoretical rate-maximizing strategy would be to either reject a patch on encounter or stay in it indefinitely (until pupation). The threshold for rejection or acceptance will depend on travel time, but patch residence time would be either zero (or minimal identification time) or lifelong. In the introduction, it appears as if the system follows the classical Marginal Value Theorem assumptions as used in classical foraging theory. In that case, patch residence time is fundamentally sensitive to a decline in intake rate while in a patch. This raises questions about what factors drive patch-leaving in the present protocol. A better theoretical framework would focus on behavioural variables that can be expected to depend on the circumstances of the experiment, as discussed below.

      (2) Rather than make predictions about time in the patch, which as explained above do not reflect the present system, larval behaviour could be modelled and described as a function of observable properties such as: (a) speed of locomotion; (b) tendency to deviate from straight progress (area restricted searching); (c) probability of return after leaving a patch, possibly controlled through rea restricted searching; (d) a response to concentration gradient, since patch boundaries are probably gradual through diffusion. There is a useful literature in this regard in studies of parasitic wasps such as Venturia canescens (formerly Nemeritis canescens, see Waage 1979). Larva may respond directly to local resource concentration (see van Alphen, J. J., Bernstein, C., & Driessen, G., 2003), where higher concentration leads to increased feeding rate, reduced locomotion, and consequently results in longer time in each patch. This could still be a normative model, but based on realistic driving inputs. The dimensions of the system make it unlikely that larvae have the opportunity to adjust to travel time, or patch composition, on which classical foraging models are based. The original versions of the marginal value theorem were thought for cases where birds exploited pine cones, so that each bird had multiple encounters, and also on dung flies that mated in dung patches, which also dried out. A system with heritable optimised parameters could work for other natural systems where the parameters can be heritable, but not here.

      (3) The previous argument indicates that patch time, while it is a real quantitative consequence, is not ideal as the major dependent variable for this system. Given that the authors have the full trajectories, they could treat movement in discrete time bins and ask if the tendency to depart from linear progression (i.e. from moving straight ahead) is a function of the density of the resource. It would appear as if all the results, including return to patches (but not memory), could be explained by area-restricted searching (see Dorfman, A., Hills, T. T., & Scharf, I. (2022). A guide to area‐restricted search: a foundational foraging behaviour. Biological Reviews, 97(6), 2076-2089.). Slower movement (perhaps directly caused by eating) and more twisted progress could generate longer times in higher food densities.

      (4) The evidence for an effect of prior experience is interesting but could be strengthened. The authors state that PRT on the second patch depends on the concentration in the first patch. However, statistically significant modulation of prior experience was only found when the second food patch was richer, namely 1M fructose (Figure 4C). If the change in patch time is due to a form of learning and contrast, one might expect significantly shorter times in any second patch if the first one was richer, which is not the case. One difficulty is that the 'patchy' nature of the environment may not be evident to the larvae, because they are much smaller than the patches. From a larva's perspective, a patch is an environment, potentially suitable to remain in until pupation (which is what they ought to do in richer food patches).

      (5) The modelling section is promising but currently somewhat underdeveloped relative to the strength of the claims. The authors fit a drift-diffusion model to data and report that a drift-only model captures homogeneous environments, whereas adding a leak term improves the fit in heterogeneous environments. This provides a useful quantitative summary of behavior but the biological interpretation of the leak parameter is not clear. In addition, the valence condition was not modelled.

    4. Reviewer #3 (Public review):

      Summary:

      The work investigates how the foraging behaviour of Drosophila larvae depends on resource quality, valence, and heterogeneity in the foraging environment. A specific focus of the work was to study how foraging decisions depend on the prior experience of alternative resource patches in the same environment. Moreover, the work presents computational models (drift diffusion models) that recapitulate foraging decisions, and whose parameters appear to depend on resource quality and environment statistics, providing potential insights into the dynamics of the decision-making process.

      I am not familiar with previous literature on foraging decisions in Drosophila, but I was specifically consulted to comment on the computational modelling. Therefore, my comments will mostly focus on the modelling aspects.

      Strengths:

      In my understanding, the two strengths of the current study are that:<br /> (1) it uses non-volatile resources, providing better control of the available cues that could guide foraging decisions, and<br /> (2) it tracks foraging behaviour over an extended period of time (3h), generating a rich dataset of foraging behaviour in the same environment.

      Overall, the study appears to have been carefully conducted.

      Weaknesses:

      The computational modelling currently provides limited additional value beyond the empirical results. There are no prior hypotheses that are addressed by the computational models. Given the flexibility of DDMs, fitting foraging times is expected to be feasible. The question is whether the fits provide mechanistic insight. The main insight appears to be that describing foraging times in a homogeneous environment requires a single free parameter (drift rate), while the heterogenous environment requires a second parameter (leak). However, the effective complexity of the model is higher than the stated parameter count suggests, as each patch quality is fit with a different drift rate, which does not generalise across environments: in the heterogeneous environment, the drift rate differs substantially across fructose concentrations, whereas in the homogeneous environment, the same concentrations yield nearly identical drift rates. Counter their claims, the authors also do not systematically explore the effect of specific prior foraging experience on computational parameters, but only contrast model fits to environments with different statistics, in which prior experiences will be generally different. Overall, at the moment these modelling results have a rather descriptive character, and provide very little insight into the underlying computational principles that drive foraging decisions.

      A second weakness is that the study does not report the detailed results of the statistical tests, and it seems that the authors interpret several differences that are not marked as statistically significant in the figures. Furthermore, the model comparisons do not account for different degrees of freedom of the models, and the goodness of fit values alone are insufficient to conclude that one model is better than the other (rather than overfitting).

    1. eLife Assessment

      This useful study investigates noise-robust and energy-efficient circuit mechanisms for working memory by optimizing connectivity and reports that the resulting networks exhibit rotational dynamics and better match aspects of PFC population recording. However, the supporting evidence remains incomplete, given the restricted linear, task-specific training and analysis, and limited comparisons with other prominent models. The manuscript would be strengthened by extending the analysis to nonlinear dynamics, providing more rigorous comparisons with alternative models, and establishing a stronger link to prior theoretical and experimental work.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors address the question of working memory maintenance, starting from the experimental observation that recordings of neural activity during the delay period of working memory tasks are sometimes observed to be dynamic. They introduce a new combination of metrics (noise-robustness and energy efficiency) to quantify the performance of various network mechanisms of memory maintenance, in linear networks. They compared attractor networks, feed-forward networks, and networks trained with a loss that includes a robustness and an energy-efficiency component. They show, by plotting state-space trajectories, that networks optimized with this loss exhibit a form of rotational dynamics. They analyzed the data recorded during the delay of a working memory task in PFC, and observed state-space trajectories similar to those of the trained networks.

      The comparison with other network mechanisms is interesting in principle, but limited by the fact that only linear networks are considered. This led to counter-intuitive and misleading statements, like the fact that attractor networks are not robust to noise, or that feed-forward networks have energy consumption that is exponential in the number of neurons.

      Strengths:

      (1) The idea to use both robustness to noise and energy efficiency to assess the performance of networks on working memory tasks is interesting.

      (2) The manuscript is clearly written.

      (3) There is an interesting combination of methodologies: theory on simple models, network training, and data analysis.

      Weaknesses:

      (1) Linear networks only.

      The main feature of attractor networks is their robustness to noise, which is typically allowed by the non-linearity of neural responses. To fit their modeling framework, the authors focused only on continuous attractor neural networks (e.g., Seung 1996) and ignored point-attractor models such as the Hopfield model, which are typically used to model WM tasks, and which would presumably lead to very different results, e.g., in Figure 1D.

      The linearity assumption is also problematic for the comparison with feed-forward models. It seems that the authors obtained runaway firing rates, explaining Figure 1F middle, which are typically prevented in non-linear networks.

      The choice of parameters for the attractor network in Figure 1 is not explained. Why is t_slow = 10^4 chosen, and what does it correspond to? We expect in linear networks that activity goes back to zero or diverges as an exponential, but in principle, the time constant can be chosen to be of the same order as the time delay, with approximately linearly decreasing SNR.

      Regarding the comparison of the different mechanisms, it would have been nice to better define the notion of rotational dynamics, beyond only considering state-space analysis, which is limited to providing mechanistic interpretations.

      (2) Fixed duration of delay periods.

      I have understood that for a given network, the duration of the delay period is fixed, as opposed to a delay duration that would fluctuate from trial to trial. This would be an important assumption to relax as well, to better match common experimental paradigms, as well as to expose a fairer comparison with other network mechanisms. See Orhan and Ma (2023) for such a discussion.

      (3) Relationship with previous works

      Many other works addressed the question of dynamic firing rates during maintenance periods of WM tasks; they should be discussed and compared to the mechanism proposed here. This includes: Barak et al, Progress in Neurobio. 2013, Pereira-Obilinovic, Aljadeff, Brunel, PRX 2023, Hansel, Mato, 2013, or works pertaining to the activity-silent neural states (allowed by short-term plasticity), the framework in which the data of Panichello et al are interpreted in the original publication.

    3. Reviewer #2 (Public review):

      In this manuscript, Ritter et al. propose a model of working memory (WM) that combines feedforward and rotational dynamics. The model is discovered by optimizing a linear RNN using a loss function that encourages maximization of signal-to-noise ratio (SNR) and minimization of activation magnitude. The authors argue that the optimized model outperforms other WM models in terms of SNR and energetic efficiency, while also better replicating key features of neural responses recorded in monkey pre-frontal cortex (PFC) during a WM task. The authors also draw connections to state space models (SSM) used for other machine learning applications.

      My main issue with this manuscript is that it does not appear to convincingly demonstrate that rotational dynamics offer any advantage over purely feedforward dynamics. The authors adopt three criteria according to which they compare models:<br /> (1) SNR.<br /> (2) Energy efficiency.<br /> (3) Similarity to neural data.

      In terms of SNR, purely feedforward models seem to perform similarly to the optimized models (Figure 1). Figure 1 does seem to show that the optimized network produces responses of smaller magnitude when the number of units is large, but the authors do not explain why adding rotational dynamics would produce such a relationship. In fact, the responses that are plotted for the feedforward network in Figures 1B, 2C, and 5E look similar, if not smaller in magnitude than those of the optimized model. Lastly, while the authors claim in the body of the text that the optimized model replicates key features of monkey PFC responses better than the purely feedforward model, this is not apparent to me from the comparisons plotted in Figure 5E-J. The authors thus do not show strong evidence that the model they propose beats what they claim is an established baseline on any of the three criteria.

      Another weakness of the manuscript is that the comparison to attractor and feedforward models seems somewhat unfair. In Figure 1, the rotational model is optimized, while the parameters for the attractor and feedforward models seem to have been at least partially chosen by hand. Figure 5C again shows the three models side by side, but the fact that it compares the same network at different stages during training complicates the comparison. Instead, one should compare the rotational solution to the optimal attractor and feedforward models, respectively (obtained by constrained optimization). From looking at the flow-fields, it seems that a feedforward network with an optimized level of amplification may work just as well. On a mechanistic level, it is unclear what computational advantage rotations offer over feedforward dynamics in the WM context.

      The choice of baseline models to compare against might be questionable. The simple line attractor model by Seung et al. (1996) was initially designed to explain oculomotor integration. It is true that a line attractor has been suggested as a mechanism for working memory, e.g., in the seminal work by Machens et al (2005). However, it seems fair to say that most studies employing non-linear networks have focused on point attractors as mechanisms of working memory (e.g., Wong & Wang, 2006; Driscoll, Shenoy, Sussillo, 2024). A point attractor arguably does not suffer the SNR issues of a line attractor, because it does not lead to integration of the noise over time. However, non-trivial point attractors cannot be implemented in linear networks of the kind studied by the authors of the present study.

      The authors should expand their discussion to include other, potentially closely related work proposing rotation-like dynamics in artificial neural networks during working memory. In particular, the manuscript does not discuss Sharma, Proca, et al, ICML 2026, which describes a rotational solution to a similar WM task obtained by optimizing linear RNNs (Sharma et al., 2026, Fig. 6). Notably, Sharma et al. arrive at a similar rotational (and likely also non-normal) mechanism without using either noisy inputs or a constraint on energy efficiency. The authors of the present manuscript should discuss to what extent this finding contradicts their claim that "normative pressures on noise-robustness and energetic cost shape the complex dynamics of WM circuits." (present manuscript, Introduction). Given the obvious parallels between the two studies, a comparison between the present work and Sharma et al. (2026) would add necessary context to the Discussion.

      The authors should also clarify the significance of the "novel method for optimization of continuous-time RNNs driven by noisy inputs" (see Discussion) that the authors propose. This method is mentioned in the first line of the Discussion section but is barely discussed, let alone sufficiently explained, in the previous Sections. The only time a comparison to BPTT with a simple MSE loss is mentioned, it is stated that the two procedures produce the same results. The novel method appears to consist of a loss with two terms, the second of which is a well-known L2-penalty on unit activations (Sussillo et al., 2015). It is not clear that the method is either novel or necessary to obtain the reported results.

      Except for the fact that higher-dimensional networks also converge on rotational solutions, Figure 3 does not add much to the reader's understanding of the optimized model (except for panel F). I find the comparison to SSMs too superficial to provide real insight.

      Figure 4 claims to show that the optimized model recapitulates "a range of properties observed in prefrontal cortex and other brain areas during WM tasks" (p. 7) but does not show neural data for comparison.

    4. Reviewer #3 (Public review):

      Summary:

      The authors optimize continuous-time linear recurrent networks driven by noisy input, computing the gradient of decoding performance numerically and analytically. Optimizing for stimulus discriminability after a delay, with a penalty on firing rate, they find networks that adopt what they call high-dimensional rotational dynamics. They argue that these outperform attractor and feedforward models on noise robustness and energetic cost, and resemble state-of-the-art state-space models. They then fit a targeted dimensionality reduction model to prefrontal recordings from monkeys performing a spatial working memory task and argue that the population structure matches the rotational solution.

      Strengths:

      The evolution of the dynamics throughout learning is a nice observation, as are the analytical calculations, although I am not sure they are new since there is a fair share of work on the learning dynamics of linear networks.

      Weakness:

      I see many weaknesses. I will classify them into five groups.

      (1) Strawman comparison and no clear definition of what is rotational. The paper is centered on comparing a trained model with two models meant to represent "attractor dynamics" and non-normal dynamics. Both are picked as the weakest member of their class.

      I use quotation marks for "attractor dynamics" because I am not sure a linear system with an eigenvalue equal to zero is a representative model for the class. This is a particular linear instantiation of the line attractor from Seung 1996, but most attractor models are nonlinear and far more robust to noise, and they are robust through error correction that this linear model does not have. Even modern continuous attractors (Rivkind and Darshan) are very robust to noise through multiple mechanisms. So what the authors picked as an "attractor model" is a limited zero-eigenvalue case that, of course, will drift. "Attractor networks are highly susceptible to noise" is therefore true only of the toy they built, not of the class.

      Second, what they call a non-normal model is in fact a feedforward chain, the extreme of non-normality. There are degrees of non-normality in any matrix, and the homogeneous delay line is the corner that requires the largest firing rates. This is not representative. See Daie et al., which has a skip and recurrent structure, or Stroud, which is not a pure chain. So the feedforward chain was also picked as a strawman, chosen so that the energetic cost they then complain about is guaranteed.

      This brings me to the real problem in this section. "Rotational" is never defined. If it means complex eigenvalues, then it is a spectral property of any non-normal matrix, and "rotational versus feedforward" is not a dichotomy; it is two regions of the same continuous space of non-normal connectivity. Their own Figure 2C shows the network passing continuously through an attractor, then feedforward, then rotational during optimization. If these are points on a continuum, then "rotational dynamics is optimal" is just a statement about where the optimizer lands under this particular loss and input normalization, not the discovery of a new dynamical class. They need to define the term operationally and show the solution is qualitatively, not just quantitatively, different from non-normal feedforward. I do not think it survives that test.

      This brings me to the references.

      (2) The dynamical mechanisms of working memory have been studied for more than two decades, and I am surprised how much directly relevant work is missing. First, Druckmann and Chklovskii 2012, where a linear system produces stable encoding from oscillating modes. This is essentially their result more than a decade earlier, and it is not cited. They also miss Murray et al. on stable encoding and heterogeneous timescales in data. They oversimplify the attractor picture; for example, Pereira-Obilinovic et al. 2023 show you can have genuinely stable attractors. They do cite Daie et al., but they ignore its central claim, that non-normality is the underlying mechanism, which is more troubling than not citing it because it means they read it and did not engage. Overall, the references are idiosyncratic, missing relevant work, and not engaging the results of papers they cite.

      This brings me to the third point.

      (3) Novelty and the relationship to Stroud and Orhan. Those papers take a similar optimization approach and find that, depending on the task parameters, the optimal solution is non-normal, non-normal plus attractor, or attractor. My impression is that what this work calls rotational is just the dynamics of a strongly non-normal A, selected here by the firing-rate regularizer. They never clarify the connection with Stroud. Is the only difference the energy penalty?

      The way to settle this is quantitative, and they have the handle and do not use it: report the Henrici departure-from-normality of their optimized A and place the solution inside Stroud's regime structure.

      There is also a tension they leave implicit. In Stroud, the early loading direction is orthogonal to the late persistent readout, and that orthogonality is the source of dynamic coding. This paper's subspace alignment result (Figure 5G, H) shows exactly this early-to-late orthogonalization in both model and data, and then presents it as evidence for the rotational account and against Stroud's hybrid. You cannot reproduce a Strout's stim vs. decoder orthogonality and claim it against Strout's without doing more work.

      (4) I did not understand the SSM section, and I think it should be cut. Is this a result? Either "SSM" just means a linear dynamical system, in which case it is trivial since every linear network here, including the LMU is an SSM, or it means the network matches a fixed-connectivity model like the LMU, which it does not seem to either. So in what sense is it a result?

      (5) The data analysis is one section, and the analysis could be described as feeling somewhat like an afterthought on a very rich dataset. The coding structure they show for the rotational model also looks like the Stroud non-normal-plus-attractor model to me. They even state that the hybrid reproduces the cross-temporal subspace. What are the quantitative, cross-session metric that discriminates rotational from the non-normal-plus-attractor hybrid? Is it eyeballed trajectories?

    1. eLife Assessment

      This important study provides a detailed characterization of individual sarcomeres' contractility and of their synchrony in spontaneously beating cardiomyocytes derived from human induced pluripotent stem cells. The combination of high-resolution tracking, statistical analysis and mesoscopic modeling leads to compelling evidence that sarcomeres operate as dynamically unstable units, leading to stochastic heterogeneities in their contraction-elongation cycles depending on substrate stiffness. The work will be relevant to scientists interested in muscle biophysics, nonlinear dynamics and synchronization phenomena in biological systems.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this manuscript, the authors present comprehensive experimental observations and a theoretical framework to explain the heterogeneous behaviour of sarcomeres in cardiomyocytes. They show that a stochastic component exists in their contractile activity, which may act as a feedback mechanism regulating physiological function.

      Strengths:

      Experiments and data analysis are robust and valid. The rigorous statistical analysis and unbiased methods enable the authors to draw well-supported conclusions that go beyond the existing literature. Their outcomes inform about cellular activity at the individual level and the authors explain how the transient dynamics of single sarcomeres are governed by a force-velocity relationship and lead to the complex contractile patterns. The similarity of the results to the study cited in [24] demonstrates the validity of the in vitro setup for answering these questions and the feasibility of such in-vitro systems to extend our knowledge of out-of-equilibrium dynamics in cardiac cells.

      Very interesting the suggestion that the interplay between intrinsic fluctuations and the dynamic instability are part of a feedback mechanism for maintaining structural and functional homeostasis.

      The addition of the theoretical model and the new text of the manuscript improves the clarity of the study.

    3. Reviewer #2 (Public review):

      Summary:

      Sarcomeres, the contractile units of skeletal and cardiac muscle, contract in a concerted fashion to power myofibril and thus muscle fiber contraction.

      Muscle fiber contraction depends on the stiffness of the elastic substrate of the cell, yet it is not known how this dependence emerges from the collective dynamics of sarcomeres. Here, the authors analyze contraction time series of individual sarcomeres using live imaging of fluorescently labeled cardiomyocytes cultured on elastic substrates of different stiffness. They find that a reduced collective contractility of muscle fibers on unphysiologically stiff substrates is partially explained by a lack of synchronization in the contraction of individual sarcomeres.

      This lack of synchronization is at least partially stochastic, consistent with the notion of a tug-of-war between sarcomeres on stiff sarcomeres. A particular irregularity of sarcomere contraction cycles is 'popping', the extension of sarcomers beyond their rest length. The statistics of 'popping' suggest that this is a purely random process.

      Strengths:

      This study thus marks an important shift of perspective from whole-cell analysis towards an understanding the collective dynamics of coupled stochastic sarcomeres.

    4. Reviewer #3 (Public review):

      The manuscript of Haertter and coworkers studied the variation of the length of a single sarcomere and the response of microfibrils made by sarcomeres of cardiomyocytes on soft gel substrates of varying stiffness.

      The measurements at the level of a single sarcomere are an important new result of this manuscript. They are done by combining the labeling of the sarcomeres z line using genetic manipulation and a sophisticated tracking program using machine learning. This single sarcomere analysis shows strong heterogeneities of the sarcomeres that can show fast oscillations not synchronized with the average behavior of the cell and what the authors call popping events which are large amplitude oscillations. Another important result is the fact that cardiomyocyte contractility decreases with the substrate stiffness, although the properties of single sarcomeres do not seem to depend on substrate stiffness.

      The authors suggest that the cardiomyocyte cell behavior is dominated by sarcomere heterogeneity. They show that the heterogeneity between sarcomere is stochastic and that the contribution of static heterogeneity (such as composition differences between sarcomeres) is small.

      Strengths:

      All the results are, to my knowledge, new and original. The authors also made a theoretical model where each sarcomere is described by a Langevin equation based on a non-linear coupling between force and velocity of the sarcomeres. This model accounts well for the experimental results including the observation of what the authors call popping events.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors present comprehensive experimental observations and a theoretical framework to explain the heterogeneous behaviour of sarcomeres in cardiomyocytes. They show that a stochastic component exists in their contractile activity, which may act as a feedback mechanism regulating physiological function.

      Strengths:

      Experiments and data analysis are robust and valid. The rigorous statistical analysis and unbiased methods enable the authors to draw well-supported conclusions that go beyond the existing literature. Their outcomes inform about cellular activity at the individual level and the authors explain how the transient dynamics of single sarcomeres are governed by a force-velocity relationship and lead to the complex contractile patterns. The similarity of the results to the study cited in [24] demonstrates the validity of the in vitro setup for answering these questions and the feasibility of such in-vitro systems to extend our knowledge of out-of-equilibrium dynamics in cardiac cells.

      Very interesting the suggestion that the interplay between intrinsic fluctuations and the dynamic instability are part of a feedback mechanism for maintaining structural and functional homeostasis.

      The addition of the theoretical model and the new text of the manuscript improves the clarity of the study.

      Reviewer #2 (Public review):

      Summary:

      Sarcomeres, the contractile units of skeletal and cardiac muscle, contract in a concerted fashion to power myofibril and thus muscle fiber contraction.

      Muscle fiber contraction depends on the stiffness of the elastic substrate of the cell, yet it is not known how this dependence emerges from the collective dynamics of sarcomeres. Here, the authors analyze contraction time series of individual sarcomeres using live imaging of fluorescently labeled cardiomyocytes cultured on elastic substrates of different stiffness. They find that a reduced collective contractility of muscle fibers on unphysiologically stiff substrates is partially explained by a lack of synchronization in the contraction of individual sarcomeres.

      This lack of synchronization is at least partially stochastic, consistent with the notion of a tug-of-war between sarcomeres on stiff sarcomeres. A particular irregularity of sarcomere contraction cycles is 'popping', the extension of sarcomers beyond their rest length. The statistics of 'popping' suggest that this is a purely random process.

      Strengths:

      This study thus marks an important shift of perspective from whole-cell analysis towards an understanding the collective dynamics of coupled, stochastic sarcomeres.

      Reviewer #3 (Public review):

      The manuscript of Haertter and coworkers studied the variation of the length of a single sarcomere and the response of microfibrils made by sarcomeres of cardiomyocytes on soft gel substrates of varying stiffness.

      The measurements at the level of a single sarcomere are an important new result of this manuscript. They are done by combining the labeling of the sarcomeres z line using genetic manipulation and a sophisticated tracking program using machine learning. This single sarcomere analysis shows strong heterogeneities of the sarcomeres that can show fast oscillations not synchronized with the average behavior of the cell and what the authors call popping eveents which are large amplitude oscillations. Another important result is the fact that cardiomyocyte contractility decreases with the substrate stiffness, although the properties of single sarcomeres do not seem to depend on substrate stiffness.

      The authors suggest that the cardiomyocyte cell behavior is dominated by sarcomere heterogeneity. They show that the heterogeneity between sarcomere is stochastic and that the contribution of static heterogeneity (such as composition differences between sarcomeres) is small.

      Strengths:

      All the results are, to my knowledge, new and original. The authors also made a theoretical model where each sarcomere is described by a Langevin equation based on a non-linear coupling between force and velocity of the sarcomeres. This model accounts well for the experimental results including the observation of what the authors call popping events.

      We thank you and the reviewers for the positive evaluation of our revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Origin of the 3-Hz oscillation and required model extension. These oscillations are reproduced by our model, and their origin is already discussed in the manuscript (see lines 403–406).

      (2) Inclusion of all 5085 LOIs vs. the selected 2321. We have expanded the explanation of the LOI selection criteria in the manuscript and clarified that the main conclusions are not sensitive to this choice (lines 161-166)

      (3) Fig. 3G caption — popping rate. The caption has been updated to clarify the units and normalization. 

      (4) Fig. 4G — "Length x" vs. ΔL. Notation corrected for consistency.

      (5) Fig. 4G — gray data points. Confirmed: these represent the mean, and the caption has been updated accordingly.

      (6) Relation of k_l to the true substrate stiffness. We have added the following clarification: "The model evaluation compared the distributions of sarcomere length changes and velocities from simulations with representative experimental LOIs from substrates (5, 15, and 85 kPa, mapped to k_l = 0.5, 1.5 and 8.5 in our 1-D model; k_l is unitless, so only the ratios between values are meaningful — rescaling k_l leaves model output unchanged under correspondingly rescaled parameters) covering the full range of mechanical loads." (lines 365-369)

      (7) Could a simpler model fit the data? The cubic polynomial in Eq. (3) was deliberately chosen as a generalist ansatz rather than imposed: its coefficients were obtained by data-driven inference via Differential Evolution, and if lower-order terms within this family had sufficed, the higher-order coefficients would have been driven toward zero. The inferred nonmonotonic force–velocity relation has two extrema separated by an unstable negative-slope branch, which sets a lower bound on the polynomial order — a linear F–v is monotonic and a quadratic admits only a single extremum, so cubic is the minimum polynomial order capable of producing the observed shape. Furthermore, the qualitative phenomena we report — popping events, dynamic instability, and stochastic heterogeneity — cannot arise from any monotonic force–velocity relation, as discussed in the section on the non-monotonic instability. With 10 parameters covering complex contractile dynamics at the individual sarcomere and myofibril level across different substrate stiffnesses, the present model is parsimonious within the family of polynomial force–velocity ansätze; we have not exhaustively searched alternative non-polynomial functional families, but any such alternative would still need to reproduce the same non-monotonic shape that the data require.

      (8) Lines 497–507 in the Discussion. On reflection, we feel these lines provide useful context for the broader interpretation and would prefer to retain them.

      (9) Line 331 — motivation of Eq. (3). We have added citations to prior work motivating this form of the equation for the broader readership.

      (10) Line 427 — "scaled". Corrected.

      Reviewer #3 (Recommendations for the authors):

      We thank the reviewer for the recommendation of a theoretical appendix. The full model code, with the formulation and implementation documented in detail, is publicly available in our GitHub repository accompanying the paper, which we believe provides a complete reference for readers wishing to explore the model further. We therefore feel an additional appendix is not necessary within the scope of this revision.

    1. eLife Assessment

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells and trachea. They show that the KO mice exhibit severe hydrocephalus due to mislocated basal bodies and impaired ciliary beating. The findings are valuable with implications in the subfield of cell biology. The evidence is solid in that the methods, data and analyses largely support the claims with only a few remaining weaknesses.

    2. Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights, regarding the function of the CCP5 N-domain.

      Comments on revised version.

      The authors have appropriately revised the manuscript in response to most of my comments.

    3. Reviewer #2 (Public review):

      Summary:

      This study analyzed consequences of Agbl5 mutation on ependymal cells development and function. Authors first characterize their mutant mouse line reporting a reduced lifespan and severe hydrocephalus. Next, they report defect in ependymal cell cilia number and motility. They provide evidence for impaired basal bodies organisation, cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicate Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype are incomplete:

      Previous comment: Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting that the author checks whether the subapical network of microtubule is glutamylated or not during ependymal cells differentiation and how this network is affected in their mutants.

      Although authors now provide images of glutamylation in figure S8 their conclusion claiming that GT335 signal is increased in cilia of Agbl5M1/M1 mutant is not supported convincingly by those pictures. Quantification would be needed.

    4. Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele by extending the deletion to the N-terminus of CCP5 to investigate its function in mouse ependymal cells and trachea.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      The manuscript is well-written, and the experiments are convincing.

      Comments on revised version.

      The authors have taken all of my comments into account and have revised their manuscript to my satisfaction.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Editors for the positive assessment on our manuscript. We also thank the Reviewers for their positive remarks and constructive comments. Based on the Reviewers’ feedback, we have conducted additional experiments and provided supporting data to address Reviewers’ comments. Particularly, we provided quantitative measurement for rotational polarity of ependymal cells in Agbl5<sup>M1/M1</sup> mutants and assessed the microtubule polarization. We quantified the intensity of apical actin network in ependymal cells to strength the role of CCP5 in organizing actin network. Using scanning electron microscopy, we demonstrated the affected polarity of trachea multicilia in Agbl5<sup>M1/M1</sup>. We co-immunostained ependymal cilia with GT335 and acetylated tubulin to address the effects on their length in cilia in the mutant. We assessed the presence and length of primary cilia in ependymal cell progenitors to identify their potential contribution to the defective polarity in Agbl5<sup>M1/M1</sup> ependymal cells. We feel that these revisions have much strengthened this MS.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights.

      We thank the Reviewer’s positive comments. To address the molecular insights of the dysregulated planar cell polarity (PCP) in Agbl5<sup>M1/M1</sup> ependyma, we have conducted additional experiments to assess the microtubule polarization in ependymal cells (Figure 7O-P). We quantified the intensity of actin networks around BB patches to better understand how it is affected in the ependyma of the mutants and contributes to the dispersion of BBs (Figure 4M-N), (Please see Recommendations for the authors).

      We also assessed trachea multicilia in Agbl5<sup>M1/M1</sup> mutants using SEM and found that the polarity of trachea multicilia was affected as well (Figure S2).

      Reviewer #2 (Public review):

      Summary:

      This study analyzed the consequences of Agbl5 mutation on ependymal cell development and function. The authors first characterize their mutant mouse line reporting a reduced lifespand and severe hydrocephalus. Next, they report a defect in ependymal cell cilia number and motility. They provide evidence for impaired basal body organisation and cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicates Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype is incomplete:

      We thank the Reviewer’s constructive comments. We have performed additional quantitative analysis of the phenotypes in Agbl5<sup>M1/M1</sup> that we feel strengthen this study.

      Figure 3G - the sequence from the movie is not really informative. Providing beating frequencies as quantification of the data would be more informative.

      We have provided the beating frequency as well as the mean vector length of cilia beating directions (that reflects the coordination of cilia) in Figure 3H and 3I respectively in the revised manuscript.

      Figure 3 - the quantification of actin network would strengthen the message.

      We agree with the Reviewers. We have quantified the total intensity of actin around BBs and the actin intensity normalized to signals of the BB marker (CEP164). The data have been provided in Figure 4M and 4N respectively. The quantitative analysis showed that both the total intensity of apical actin network and the intensity of F-actin per BB are reduced in Agbl5<sup>M1/M1</sup> ependymal cells compared to that in wild-type mice, suggesting that CCP5 is involved in organizing actin network around BB. This analysis certainly improves the clarity of this message.

      Lines 219 -220 - the authors conclude «Taken together, in Agbl5M1/M1 ependymal cells, the expression of genes promoting multiciliogenesis were not impaired but certain proteins associated with differentiated ependymal cells are not properly expressed». However, they do not assess gene but protein expression (IF). In addition, their quantification shows differences in the number of FoxJ1 positive cells which indeed is an impaired expression.

      We will clarify this statement and emphasize the number of FoxJ1-positive cells.

      Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting for the authors to check whether the subapical network of microtubules is glutamylated or not during ependymal cell differentiation and how this network is affected in their mutants.

      We thank the Reviewer’s constructive comments. We conducted an immunostaining on whole-mount lateral walls of lateral ventricles for GT335 and Centrin1, the position of the latter being used to localize the subapical layer. While the GT335 signal in multicilia is increased in Agbl5<sup>M1/M1</sup> ependyma (Figure S8E), its signals underneath BBs are not much different between the mutant and wild-type (Please see Figure S8C, D, G, H).

      Showing the data mentioned in the discussion on Cep110 would be a nice addition to the paper.

      These data have been provided in Supplementary Figure S9.

      Line 354: "The latter serves as a component of tissue polarity that is required for asymmetric PCP protein localization in each cell (Boutin et al., 2014; Vladar et al., 2012)." The cited reference did not demonstrate that this microtubule network is required for asymmetric PCP localization.

      We thank the Reviewer for critical reading. The cited reference (Bountin et al., 2014) has been removed.

      Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      Weaknesses:

      The manuscript is well-written but lacks specific interpretations of the results presented. Further experiments are needed to be fully convincing.

      We thank the Reviewer’s comments. We have performed further analysis and conducted additional experiments to strengthen this study.

      (1) We have quantified the intensity of actin staining around BB patches and its intensity relative to the number of BBs to assess to which extent the actin networks in Agbl5<sup>M1/M1</sup> ependymal cells are affected (please refer to the above response to the comments of Reviewer 2#). The results were shown in Figure 4M-N.

      (2) We Co-stained tdTomato with an ependymal cell-specific markers to strengthen the expression of Agbl5 in ependymal cells (please see Figure 6C-E).

      (3) We have conducted co-immunostaining of GT335 and Ac-Tub and compared the length of their signals in ependymal multicilia between WT and Agbl5<sup>M1/M1</sup> mice (please see Figure 6O, P, R, S).

      (4) We quantified the area of ependymal cells in the wild-type and Agbl5<sup>M1/M1</sup> mice. Indeed, the area of ependymal cells is increased in the mutants. However, the primary cilia are present in the ependymal cell progenitors of Agbl5<sup>M1/M1</sup> mice and have similar length with that in the wild-type (Please see Figure 7M, N and our response to this point below).

      (5) We performed additional analysis to address the affected rotational polarity in the Agbl5<sup>M1/M1</sup> mutant mice (please see Figure 3I, Figure 7E).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors showed that the actin networks were severely affected, leading to impaired stability of basal bodies and that the intensity and length of acetylated tubulin signal in the multicilia were dramatically reduced in AGBL5M1/M1mutant mice (Figures 3 and 5). Data also suggested the dysregulation of planar cell polarity. Are expression and localization of other planar cell polarity proteins such as tyrosinated tubulin and Fzd6 affected in mutant mice?

      We thank the Reviewer’s recommendations. We have assessed the expression of tyrosinated tubulins and found they are similarly polarized in ependymal cells from wild-type and Agbl5<sup>M1/M1</sup> mice. The results are presented in Figure 7O, P in the revised MS. We also tried to assess the expression of Fzd6. However, with the antibody we tested, Fzd6 signals were not convincing. Therefore, we prefer to not showing the results and drawing a conclusion on it.

      (2) The phenotype of multiciliated cells in tracheas should also be examined in mutant mice. It is important to elucidate whether AGBL5 commonly functions in multiciliated cells of other organs.

      We thank the Reviewer’s suggestion. We have assessed the multicilia in the tracheas of P30 mice using scanning electron microscopy. Indeed, unlike the multicilia in wild-type mice that orientate to the same direction, those in the tracheas of Agbl5<sup>M1/M1</sup> mice often radiate to different directions in individual cells (Figure S2). Therefore, Agbl5 appears commonly involved in the alignment of multicilia.

      (3) According to Figure 1B, AGBL5 is highly expressed in the brain. Which cells in the brain express it besides ependymal cells?

      Based on the localization of tdTomato tracer engineered in Agbl5 mutant alleles (Figure 5B), Agbl5 is broadly expressed in the brain, including most if not all neurons, but its expression is much weaker in the subventricular zone (Please see Figure 5B). We clarified this in the revised MS.

      (4) From a mechanistic point of view, it is necessary to identify binding proteins with the N-domain of AGBL5 and perform functional analyses.

      We agree with the Reviewer. We feel that identification of the binding partners of CCP5 N-domain and functional analysis may be more suitable to go along with other mechanistic analysis on the function of CCP5 in ependymal cell polarities in our future study.

      Reviewer #2 (Recommendations for the authors):

      (1) Movie 3: The authors could comment on beating direction that seems impaired at the cell scale here, analysis of rotational polarity would be a plus.

      We thank the reviewer’s recommendation. We have analyzed the beating directions of cilia in individual cells and presented their consistency in each cell using mean vector length. These results indeed demonstrated defective rotational polarity in the cell level in Agbl5<sup>M1/M1</sup> mice (please refer to Figure 3I). We also analyzed the beating directions of ependymal multicilia in earlier stage in tissue level (Figure 7E). The mean vector length of cilia beating direction in Agbl5<sup>M1/M1</sup> mice is significantly reduced compared to that in wild-type, suggesting an aberrant rotational polarity in the tissue level in the mutant (Figure 7E).

      (2) Line 166 : ref to Werner et al., 2011 is not correct (no ependymal cells in that paper).

      We thank the reviewer’s critical reading. This reference has been removed.

      (3) Figure S4: B and D look similar picture to me same for C and F.

      We apologize for using the wrong images in this Figure. It has been corrected (Revised Figure S5).

      (4) Line 328: "Therefore, CCP5 apparently contributes to the establishment of both translational and tissue polarities in ependymal cells." Should be rephrased since translational polarity is also a tissue-level parameter which is the coordinated positioning of the ciliary patch. Cf Mirzadeh et al., 2010; Boutin et al., 2014.

      We thank the Reviewer’s comments. The sentence has been rephrased. This concept has been clarified where else needed in the revised manuscript. 

      (5) Line 348: "Planar cell polarity (PCP) pathway is essential for the establishment of rotational and tissue polarities in ependymal cells" Rotational polarity also has a tissular component (ie coordination of beating direction across tissue which is reflected by coordination of basal body polarities across tissue).

      We thank the Reviewer’s comments. We have clarified this point in the revised MS.

      (6) Incomplete bibliography citation (ie Walentek et al. without date).

      We thank the Reviewer’s critical reading. This bibliography citation has been fixed.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 3: The authors assert that the mutant's apical actin networks are significantly disrupted. However, the cell shown in Figure 3Q-R exhibits less compact centrioles than the controls, which could account for the reduction in phalloidin staining. Because centriole dispersion is variable in the mutant, quantifying actin staining in representative cells would be necessary to support such a statement.

      We thank the Reviewer’s comments. To address this concern, we have quantified the total intensity of actin network around BBs as well as the intensity of F-actin signals normalized to the level of immunosignals of BBs ((revised Figure 4M, N) please also refer to our response to Reviewer 1#). The results indicated the intensity of actin signal per BB is reduced in the mutant compared to that of wild-type mice. We feel that this analysis strengthened our statement.

      (2) Figures S3 and 4A-B show that the authors examine tdT expression to show that Agbl5 is expressed in ependymal cells but not in the SVZ. However, the tdT signal intensity is very low, and cells are very dense in this brain region. Double staining with specific markers of ependymal and/or SVZ cells would help convince readers that tdT is not expressed in SVZ cells.

      We agree with the Reviewer that the intensity of tdT signal is low, but broadly detectable in brain. Compared with its expression in ependymal cells, that in SVZ is much lower if any (Figure 4B’). To further confirm the identity of tdT-positive cells along the surface of ventricles, we have co-stained the brain sections of Agbl5<sup>WT/M1</sup> mice for tdT and S100b, a marker of mature ependymal cells (Figure 5C-E). The signal of tdt is colocalized with that of S100b and is much lower in cell layers next to S100b-positive cells.

      (3) Figure 4C-D and S4: The authors demonstrate that the number of FoxJ1+ cells per section increases at P7 (4C-E), while the number of S100β+ cells per mm decreases. Quantifications should be carried out in a similar manner to ensure comparability (number of positive cells per mm). Additionally, it remains unclear how to interpret these results, as S100β and FoxJ1 are two markers of differentiated cells, yet they exhibit opposite trends compared to controls. Is this a direct or indirect effect of Agbl5 mutation? The increase in the number of FoxJ1+ cells is particularly surprising given that the number of GT335 multicilia per mm remains unchanged (Figure 5).

      We agree with the Reviewer that quantifications should be carried out in a similar manner. In the revised MS, the quantification of Foxj1-positive cells is presented in number per mm (Figure 5I). To be noted, the expression of Foxj1 was assessed at P7 when ependymal cells are differentiating. while the expression of S100β was assessed at P17 when ependymal cells are supposed to be fully mature. Although S100b is used as a marker of mature ependymal cells, given its unclear function, we removed the results of S100b-positiving cell counting to avoid confusion in the revised manuscript.

      (4) Figure 5: In this figure, the authors analyze the labeling obtained with GT335, Acetylated Tubulin, and Arl13b antibodies. They show that the area of the cilium labeled by GT335 has increased, while the area labeled by the Acetylated Tubulin antibody has decreased in the knockout (KO) compared to the control. However, the length of the cilia observed through labeling with the Arl13b antibody remains unchanged. These observations are intriguing, but the low-magnification images in Figure 4 do not allow for the differences in ciliary axoneme labeling to be seen. Double GT335/AcTub labeling and higher magnifications are necessary for improved visualization of the differences in labeling along the axonemes.

      We thank the Reviewer comments. We have co-stained the cilia with GT335 and Ac-Tub antibodies, re-quantified cilia length labeled with respective antibodies and provided high magnification images. Please see the revised Figure 6O,P,R,S.

      (5) Figure 6: An analysis of ciliary beats using a high-speed camera shows no difference in ciliary beat frequency between the control and KO groups. At least, 3 animals should be analyzed. According to Figure 5, these findings indicate that the decrease in ciliary acetylation and the increase in ciliary glutamylation do not affect the beat frequency; instead, they disrupt the orientation of the beats. While these results are intriguing, they require further confirmation. Analyzing ciliary beats with a high-speed camera is informative, but at least three animals per genotype should be examined to ensure rigor. Furthermore, if the coordination of ciliary beats is impaired within the cells, this should be validated by double-labeling centrioles and basal feet to demonstrate that the orientation of cilia within the cells is abnormal.

      We thank the Reviewer’s comments. Sections shown in Figure 5 (currently Figure 6) are from P7 mice, while the ciliary beating analysis shown in Figure 6 (currently Figure 7) is from P15 mice. As the PTM changes in cilia were also observed in Agbl5<sup>M2/M2</sup>, we don’t think this is the cause that disrupts the orientation of the beats. The rotational polarity of Agbl5<sup>M1/M1</sup> ependymal cells is affected. Please refer to the analysis in Figure 3I and Figure 7E in the revised manuscript.

      (6) Figure 6F-G: β-Catenin labeling reveals cells of varying sizes in the KO. This phenotype is typical of ciliary mutants that lack primary cilia (Mirzadeh et al., 2010). Hence, it is essential to examine the mutation's impact on the presence, length, and positioning of the primary cilium in ependymal cell progenitors.

      We thank the Reviewer’s constructive comments. We assessed the area of ependymal cells labeled with β-Catenin. Indeed, the ependymal cells in the mutant showed larger area than that of wild-type. The ratio of the area of BB patch over that of cell surface is reduced (please see Figure 7O, P in the revised manuscript). However, primary cilia are present in ependymal cell progenitors in the mutant and exhibit comparable length with those in the wild-type (Figure S8). Due to some technique problems, we were unable to get convincing results from whole-mount ventricle walls for the primary cilium positioning at this time. We speculate that the localization of certain sensory proteins in primary cilia or the positioning of primary cilia might be affected in Agbl5<sup>M1/M1</sup> mice. We discussed this possibility and will certainly systemically assess this intriguing aspect in our future investigation.

      (7) Given the regular beating frequency in the KO at P15, how do the authors explain the complete absence of ciliary beating in the adult? How many animals were analyzed? One would expect ciliary beating to remain unaffected as it was at P15 unless the cilia structure was specifically altered at the adult stage. Is that the case?

      We thank the Reviewer’s critical questions. We do think that the ciliary structure of Agbl5<sup>M1/M1</sup> ependymal cells is likely altered during aging. Given that only Agbl5<sup>M1/M1</sup> but not Agbl5<sup>M2/M2</sup> mice develop hydrocephalus, we speculate the N-domain of CCP5 may contribute to the integrity of ependymal multicilia. We have added this in the Discussion section. For each genotype, 2 mice were analyzed.

      (8) Line 264 of the manuscript: replace intercellular with intracellular.

      It has been revised.

      (9) Indicate the number of animals analyzed in each experiment

      It has been included in figure legends.

    1. eLife Assessment

      This paper addresses a valuable research question on the modest heritability of the brain's response to movie watching, and how heritability varies under different parameters such as regional spatial hyperalignment and BOLD frequency bands. The topic of this paper is of interest to fMRI methodological experts, and potentially to a broader cognitive neuroscience audience, and those with an interest in understanding the heritable sources of individual differences in brain function. Although some of the conclusions could be strengthened by future cross validation studies in independent and larger family-based samples, and through complementary twin/family and SNP-based models, taken altogether, the analyses and results provide convincing evidence for the overall conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales, and more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor our topographic differences and found the relationship between heritability and neural time scales very interesting. The writing is clear and the results are compelling. In general, I don't have many complaints after a couple reads through the manuscript; most of my comments below are relatively minor suggestions and points of clarification.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified:

      On page 16, you compare heritability in functional connectivity (FC) and response time series and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and time-series differences cannot, by definition). This makes me wonder how this connectivity result would change if you used intersubject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript-but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. You could even color the data points according to networks like in Figure 3C. (You also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      On page 9, if I understand correctly, you regress the vector of ISC values across parcels out of the vector of heritability values across parcels and then plot the residual heritability values. Do you center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can you explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could you go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?)

      On page 4 (line 155), you say "we shuffled dyad labels"-is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure your approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned at line 189?

      I found panel A in Figure 4 to be a little bit misleading because your parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels you're performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248-259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Comments on revised version.

      The authors have adequately addressed my previous comments. This is a strong contribution: the methods are sophisticated, the statistical treatment is rigorous, and the results are quite interesting/compelling. I'm happy to endorse the revised manuscript as a finalized version.

      Just to confirm: The subjects watched all different movies across the two days, right? For a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

    3. Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      Comments on revised version.

      The whole manuscript has been improved a lot, and the concerns have been clarified.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained by local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales and is more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor out topographic differences, and I found the relationship between heritability and neural time scales very interesting. The writing is clear, and the results are compelling.

      We thank Reviewer 1 for their kind words and enthusiastic support of our manuscript.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified.

      (1) On page 16, the authors compare heritability in functional connectivity (FC) and response time series, and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and timeseries differences cannot, by definition). This makes me wonder how this connectivity result would change if the authors used inter-subject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript, but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      We agree that investigating the heritability of ISFC (or stimulus-driven functional connectivity) would make for a very interesting future direction. Ultimately, we chose to analyze FC (vs. ISFC) profiles to allow for direct comparison with the sizable existing literature on the heritability of FC (such as in our Movie vs. Rest FC analysis) and decided to refrain from analyzing ISFC data in order to keep the present manuscript focused. ISFC analysis of this dataset will be a focus of future work.

      (2) The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. The authors could even color the data points according to networks, like in Figure 3C. (They also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      We thank R1 for this helpful suggestion- we originally set the y-axis limits to r = 1 in order to facilitate comparison between ISC (Fig. 1B) and FC profile (Fig. 6B) similarity, but we agree that this renders the group differences harder to discern and have updated the plot accordingly (along with thicker lines to enhance readability). We prefer to keep the line plots in the main body as they allow for direct comparison of all three groups on the same plot, but we have included the scatter plot version in Fig. S2 for those who are interested.

      (3) On page 9, if I understand correctly, the authors regress the vector of ISC values across parcels out of the vector of heritability values across parcels, and then plot the residual heritability values. Do they center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can the authors explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could they go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?).

      We indeed included an intercept in this model using MATLAB’s fitlm function. This means that the model estimates the best-fitting line of the following form: heritability<sub>i</sub>=β0+β1ISC<sub>i</sub> +ε<sub>i</sub>. We agree that the interpretation of these ε<sub>i</sub> values and alternative approaches to controlling for ISC should be clarified. As such, we have added the following passages to the text:

      Methods: “Because the heritability of ISC is constrained by the degree of synchronization in a given area, we also sought to identify areas in which BOLD time courses were more/less heritable than would be expected based on ISC alone by fitting a linear model of the form heritability<sub>i</sub>=β0+β1ISC<sub>i</sub>+ε<sub>i</sub> and plotting the residuals. Regarding alternative approaches to controlling for ISC, although the heritability model introduced by Ge et al. allows for the inclusion of covariates defined at the subject level (e.g., age), it does not allow for covariates that are defined at the dyad level (e.g., pairwise ISC).”

      Results: “Here, negative values in the residual map indicate parcels where heritability is lower than expected based on ISC, while positive values indicate higher-than expected heritability.”

      (4) On page 4 (line 155), the authors say "we shuffled dyad labels"- is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure their approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned in line 189?

      Briefly, shuffling the kinship matrix involved permuting the rows and columns of the matrix in the same manner (also known as the quadratic assignment procedure), whereas shuffling the dyad labels involved random permutations of the three group labels (MZ, DZ, unrelated), which could not be done through matrix operations as the age- and gender matching precluded the use of a complete similarity matrix. However, given concerns raised by Reviewer 2, we have removed our significance claims from this (and similar) sections, which we discuss in more detail in response to Reviewer 2’s weakness A.

      (5) I found panel A in Figure 4 to be a little bit misleading because their parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels they are performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      We agree that our efforts to present a simplified depiction of hyperalignment may mislead less familiar readers and have amended Fig. 4A according to this suggestion. We have also added text to the methods section (below) to clarify that the outputs of hyperalignment are time series that reflect linear combinations of other voxels’ time series from that parcel.

      “This approach independently transforms each subject's data within discrete anatomical parcels into the common space, yielding functionally aligned vertex time series that are calculated as weighted linear combinations of the original time series from all other vertices within that same parcel for that subject.”

      (6) I believe the subjects watched all different movies across the two days, however, for a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

      We agree that this would be helpful and have added the following text to the relevant sections:

      “All clips were only viewed once by each subject, with the exception of the brief montage which was included at the end of each of the four runs for test-retest purposes.”

      “To characterize the heritability of brain responses to complex stimuli, we used 7T fMRI data from 178 HCP Young Adult subjects acquired across two days (using two largely non-overlapping sets of movie stimuli, see Methods)…”

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to estimate the heritability of brain activity evoked from a naturalistic fMRI paradigm. No new data were collected; the authors analyzed the publicly available and well-known data from the Human Connectome Project. The paper has 3 main pieces, as described in the Abstract:

      (1) Heritability of movie-evoked brain activity and connectivity patterns across the cortex.

      (2) Decomposition of this heritability into genetic similarity in "where" vs. "how" sensory information is processed.

      (3) Heritability of brain activity patterns, as partially explained by the heritability of neural timescales.

      Strengths:

      The authors investigate a very relevant topic that concerns how heritable patterns of brain activity among individuals subjected to the same kind of naturalistic stimulation are. Notably, the authors complement their analysis of movie-watching data with resting-state data.

      Weaknesses:

      The paper has numerous problems, most of which stem from the statistical analyses. I also note the lack of mapping between the subsections within the Methods section and the subsections within the Results section. We can only assess results after understanding and confirming the methods are valid; here, however, Methods and Results, as written, are not aligned, so we can't always be sure which results are coming from which analysis.

      (A) Intersubject correlation (ISC) (section that starts from line 143): "We used nonparametric permutation testing to quantify average differences in ISC for each parcel in the Schaefer 400 atlas for each day of data collection across three groups: MZ dyads, DZ dyads, and unrelated (UR) dyads, where all UR dyads were matched for gender and age in years." ... "some participants contributed to ISC values for multiple dyads (thus violating independence assumptions)"

      This is an indirect attempt to demonstrate heritability. And it's also incorrect since, as the authors themselves point out, some subjects contribute to more than one dyad.

      Permutation tests don't quantify "average differences", they provide a measure of evidence about whether differences observed are sufficient to reject a hypothesis of no difference.

      Matching subjects is also incorrect as it artificially alters the sample; covarying for age and sex, as done in standard analyses of heritability, would have been appropriate.

      It isn't clear why the authors went through the trouble of implementing their own nonparametric test if HCP recommends using PALM, which already contains the validated and documented methods for permutation tests developed precisely for HCP data.

      The results from this analysis, in their current form, are likely incorrect.

      We appreciate that permutation tests do not quantify average differences and intended to write “We used non-parametric permutation testing to quantify [the significance of] average differences…”. Our intention with this analysis was not to demonstrate heritability, but rather to quantify group differences in ISC in a manner that is interpretable for readers who are unfamiliar with h<sup>2</sup> (e.g., “identical twins’ BOLD time courses were 59% more similar than those from pairs of unrelated individuals”) and motivate the formal heritability analysis used later in the paper. Indeed, all of the heritability analyses in this paper leveraged a validated multidimensional heritability method first introduced by Ge et al. (2016) and used by many other investigators since then. Furthermore, we covaried for age and sex at the subject level in all our heritability analyses, and always tested the significance of these heritability values using a validated permutation procedure (the quadratic assignment procedure; Hubert & Schultz, 1976) that respects the non-independence of dyadic data.

      Regarding the shuffling procedure used for Figure 1, while PALM is the standard for univariate, subject-level GLMs in the HCP pipeline and can accommodate nested designs (i.e., subjects within families), it is not designed to handle the unique relational dependencies of dyadic ISC analysis (i.e., the same subject contributing to multiple dyads). Although the element-wise resampling approach was the most appropriate approach available, it is known to inflate the false positive rate (Chen et al., 2016; doi:10.1016/j.neuroimage.2016.05.023); given that this analysis was simply meant to motivate our later hypothesis testing heritability analyses, we have removed significance claims from this section of the manuscript. Still, we emphasize that this has no bearing on the validity of our conclusions which were supported by our formal heritability analyses; throughout our paper we have correctly used the appropriate methods to back the stated claims.

      (B) Functional connectivity (FC) (section that starts from line 159): Here the authors compute two 400x400 FC matrix for each subject, one for rest, one for movie-watching, then correlate the correlations within each dyad, then compared the average correlation of correlations for MZ, DZ, and UR. In addition to the same problems as the previous analysis, here it is not clear what is meant by "averaging correlations [...] within a network combination". What is a "network combination"? Further, to average correlations, they need to be r-to-z transformed first. As with the above, the results from this analysis in its current form are likely incorrect.

      We regret that R2 had difficulty understanding our analysis and have added the following text to the relevant Methods section to clarify our approach:

      “For example, there are 16 parcels in the Kong et al. Auditory network and 17 parcels in the Language network, so the FC profile for a given subject’s Auditory-Language network combination consists of the (16 * 17 =) 272 correlation coefficients between all unique pairs of one parcel from each network.”

      As we stated in the previous Methods paragraph, “All Pearson r values in this and all other analyses were Fisher z-transformed before averaging (and converted back to Pearson r for visualization)”. Thus, contrary to the reviewer’s assertion, these analyses were performed correctly. Once again, we emphasize that this analysis was not intended to demonstrate heritability, but rather to describe group differences in FC in familiar units.

      (C) ISC and FC profile heritability analyses (section that starts from line 175): Here, the authors use first a valid method remarkably similar to the old Haseman-Elston approach to compute heritability, complemented by a permutation test. That is fine. But then they proceed with two novel, ill-described, and likely invalid methods to (1) "compare the heritability of movie and rest FC profiles" and (2) to "determine the sample size necessary for stable multidimensional heritability results". For (1), they permute, seemingly under the alternative, rest and movie-watching timeseries, and (2), by dropping subjects and estimating changes in the distribution.

      The (1) might be correct, but there are items that are not clearly described, so the reader cannot be sure of what was done. What are the "153 unique network combinations"? Why do the authors separate by day here, whereas the previous analyses concatenated both days? Were the correlations r-to-z transformed before averaging?

      The (2) is also not well described, and in any case, power can be computed analytically; it isn't clear why the authors needed to resort to this ad hoc approach, the validity of which is unknown. If the issue is the possibility that the multidimensional phenotypic correlation matrix is rank-deficient, it suffices that there are more independent measurements per subject than the number of subjects.

      Regarding (1), we have clarified in section 2.6 that the 153 unique network combinations reflect each unique pair of 17 Kong networks. All of our analyses, including this one, were performed separately for each day of data collection, as we state throughout the paper and visualize in our figures (although we acknowledge that, on some occasions, we [conservatively] performed FDR-correction on a combined set of p-values, as discussed in our response to K). Given that the null hypothesis for this analysis is that rest FC and movie FC are equally heritable, we are not sure why permuting rest and movie FC matrices would be invalid. All Pearson r values were z-transformed before averaging, as we stated in our paper.

      Regarding (2), we included this analysis in response to editorial concerns that our heritability analyses were not sufficiently powered, and we chose this approach because it serves as a simple way to demonstrate the stability of our results at various sample sizes whose validity is self-evident. Furthermore, this sort of subsampling approach has been used many times before in our field (e.g., Marek et al., 2022) and others (e.g., Manyara et al., 2024) to demonstrate the sample-size dependence and stability of statistical effects. We have added text explaining this to the relevant Methods section (2.6).

      (D) Frequency-dependent ISC heritability analysis (from line 216): Here, the authors decompose the timeseries into frequency bands, then repeat earlier analyses, thus bringing here the same earlier problems and questions of non-exchangability in the permutations given the dyads pattern, r-z transforms, and sex/age covariates.

      We did not use dyadic permutation testing for any of the frequency-dependent ISC analyses; rather, we used the jackknife SEMs to compare heritability across frequency bands and have added an explicit description of this to section 2.7. We have addressed the r-z transform and covariate concerns in previous comments.

      (E) FC strength heritability analysis (from line 236): Here, the authors use the univariate FC to compute heritability using valid and well-established methods as implemented in SOLAR. There is no "linkage" being done here (thus, the statement in line 238 is incorrect in this application. SOLAR already produces SEs, so it's unclear why the authors went out of their way to obtain jackknife estimates. If the issue is non-normality, I note that the assumption of normality is present already at the stage in which parameters themselves are estimated, not just the standard errors; for non-normal data, a rank-based inversenormal transformation could have been used. Moreover, typically, r-to-z transformed values tend to be fairly normally distributed. So, while the heritabilities might be correct, the standard errors may not be (the authors don't demonstrate that their jackknife SE estimator is valid). The comparison of h2 between dyads raises the same questions about permutations, age/sex covariates, and r-z transforms as above.

      We used jackknife SEs for these analyses to maintain consistency with the multidimensional heritability package used here, which only outputs jackknife SEs. We note that this jackknife approach (and the corresponding multidimensional heritability analysis) was detailed in prior work (Anderson et al., 2021), and that the leave-one-family-out jackknife has a long history of being used to estimate SEs in heritability studies, especially when working with smaller samples (Knapp et al., 1989). We are also not sure what “the comparison of h2 between dyads” means- heritability cannot be compared “between” dyads; rather, it is defined across dyads.

      (F) Hyperalignment (from line 245): It isn't clear at this point in the manuscript in what way hyperalignment would help to decompose heritability in "where vs. how" (from the Abstract). That information and references are only described much later, from around line 459. The description itself provides no references, and one cannot even try to reproduce what is described here in the Methods section. Regardless, it isn't entirely clear why this analysis was done: by matching functional areas, all heritabilities are going to be reduced because there will be less variance between subjects. Perhaps studying the parameters that drive the alignment (akin to what is done in tensor-based and deformation-based morphometry) could have been more informative. Plus, the alignment process itself may introduce errors, which could also reduce heritability. This could be an alternative explanation for the reduced heritability after hyperalignment and should be discussed. An investigation of hyperaligment parameters, their heritability, and their co-heritability with the BOLD-phenotypes can inform on this.

      To help set up our hyperalignment analyses, we have added text to the introduction explaining how hyperalignment would help to decompose heritability. The description in the Methods section included a reference to Bazeille et al., 2021, in which the hyperalignment method used here is discussed in detail. Still, we have added citations to additional papers (also cited in the Bazeille et al. paper, and elsewhere in our paper) in case that might be helpful. We note that it is not the case that all heritabilities were reduced by hyperalignment- as can be seen in Figs. 4D, 8A, and S15, hyperalignment did increase heritability in some voxels and network combinations. This would be expected under the alternative (albeit unlikely) hypothesis that functional topographies are not heritable, such that topographic variation between related individuals would obscure similarities in their (heritable) topography-independent brain responses. Recognizing that this alternative is unlikely, we believe the main novelty of this analysis comes from the magnitude of the hyperalignment effect (up to 40% of brain-wide heritability) and its spatial pattern (e.g., larger heritability decreases in visual vs. auditory cortex, the opposite of our NT result).

      We agree that we would see lower post-hyperalignment heritability if the alignment process itself introduced errors/noise, but this would be deeply surprising as hyperalignment increases ISC by design (and errors/noise could only decrease ISC). To demonstrate this, we have added Figure S7 which shows that (as expected) ISC across all voxels and subject pairs increases after hyperalignment (and that this increase is larger when hyperalignment is performed in larger parcels). Given that hyperalignment increased ISC, and that it is blind to twin status, we are unsure how it could have introduced errors that would have confounded this result.

      (G) Relationships between parcel area and heritability (from line 270): As under F), how much the results are distorted likely depends on the accuracy of the alignment, and the error variance (vs heritable variance) introduced by this.

      We agree that alignment accuracy could potentially impact parcel-level differences in how much heritability changes following hyperalignment, and we included the frequency dependent h<sup>2</sup><sub>residuals</sub> (controlling for differences in ISC) in Fig. 3 for this reason, as more accurate hyperalignment should result in greater increases in ISC, raising the heritability ceiling. We note that we observe similar relationships between parcel rank and frequency dependent changes in these residualized maps, suggesting that our parcel-level differences are not simply the result of better alignment in more sensory parcels.

      (H) Neural timescale analyses (from line 280): Here, a valid phenotype (NT) is assessed with statistical methods with the same limitations as those previously (exchangability of dyads, age/sex covariates, and r-z transforms). NT values are combined across space and used as covariates in "some multivariate analyses". As a reader, I really wanted to see the results related to NT, something as simple as its heritability, but these aren't clearly shown, only differences between types of dyads.

      We have addressed the exchangeability, covariates, and r-z transform comments above (in A). As we explained for our FC strength analyses, we are underpowered to evaluate the heritability of unidimensional traits (like the heritability of NT magnitude), and the heritability of a closely-related measure (BOLD turnover magnitude) has already been established in a larger sample of HCP subjects (https://doi.org/10.1152/jn.00402.2022). Still, we agree that more results related to the heritability of NTs would be of interest to our readers. As such, we have added an analysis in section 3.4 quantifying the heritability of multivariate NT topographies and used SOLAR to quantify the heritability of NT magnitudes, with the disclaimer that this and similar analyses are underpowered (hence the large difference in day 1 and day 2 heritability effect sizes). We also removed significance claims for the dyadic NT similarity analysis.

      (I) Significance testing for autocorrelated brain maps and FC matrices (from line 310): Here, the authors suddenly bring up something entirely different: reliability of heritability maps, and then never return to the topic of reliability again. As a reader, I find this confusing. In any case, analyses with BrainSMASH with well-behaved, normally distributed data are ok. Whether their data is well behaved or whether they ensured that the data would be well behaved so that BrainSMASH is valid is not described. As to why Spearman correlations are needed here, Mantel tests, or whether the 1000 "surrogate" maps are valid realizations of the data under the null, remains undemonstrated.

      We brought up reliability in this section because we show the reliability of our results across the two days of data collection several times in the paper. R2 is correct to point out that BrainSMASH was validated using normally distributed brain maps, and although some of our brain maps contain normally distributed values, others are right skewed (due largely to the fact that many voxels/parcels exhibit low ISC while visual/auditory areas have very high ISC). In preparing our original manuscript, we visualized BrainSMASH’s variogram outputs for one of the most skewed inputs (vertex-wise BOLD time course heritability) and found that the autocorrelation structures of the empirical and null maps were well-matched. We did not include this in the original manuscript as it is not commonplace in the field to report the variograms, see Author response image 1. Furthermore, our use of Spearman (vs. Pearson) correlations renders these distributional differences less relevant, as the Spearman correlation transforms all inputs to a uniform distribution. To empirically check that these distributional differences do not bias our results, we retested the significance of all brain map associations using the spin test (10.1016/j.neuroimage.2018.05.070), an alternative method that does not assume normally distributed inputs, and obtained identical p-values for all analyses (P<.001 in all cases).

      Author response image 1.

      (J) Global signal was removed, and the authors do not acknowledge that this could be a limitation in their analyses, nor offer a side analysis in which the global signal is preserved.

      Although we agree that GSR is a contentious preprocessing step for certain analyses, it has explicitly been shown to increase ISC signal-to-noise without compromising FC fingerprints (Graff et al., 10.1016/j.dcn.2022.101087), and it is uncommon to perform ISC analyses with and without GSR. Still, we have added additional text to our Methods section explaining our rationale for using GSR and that this could affect our results. We also re-ran our main analysis (BOLD time course heritability) with and without GSR and found that GSR had little impact on our results; we have included this in our manuscript as Fig. S4.

      Specifically, we see that GSR resulted in a slight increase in heritability (average Day 1 h<sup>2</sup> with/without GSR = .064/.060; Day 2: .068/.061) and almost no effect on the spatial pattern of our results (With GSR/without GSR Spearman ρ = .99, P<sub>brainSMASH</sub> < .001 on both Day 1 and Day 2).

      (K) FDR is used to control the error rate, but in many cases, as it's applied to multiple sets of p-values, the amount of false discoveries is only controlled across all tests, but not within each set. The number of errors within any set remains unknown.

      We agree that the FDR usage in our original manuscript was inconsistent, in that for two analyses we FDR-corrected p-values from the two days of data collection together (instead of correcting p-values from each day separately and reporting voxels/parcels/etc. that were significant at q<.05 on both days, as in the rest of our analyses). We note that both approaches are more conservative than reporting significant results at q<.05 separately; regardless, to maintain consistency we have updated all analyses such that FDR correction is always performed separately for each day of data collection.

      (L) Generally, when studying the heritability of a trait, the trait must be defined first. Here, multiple traits are investigated, but are never rigorously defined. Worse, the trait being analyzed changes at every turn.

      Here, we analyze the heritability of movie-evoked BOLD time courses (Figures 1-5) as well as FC profiles (Figures 6-8). We defined FC profiles in our Introduction as an individual’s pattern of pairwise FC strengths (and further detailed how we quantified FC profiles in the relevant Methods section), and believe that “BOLD time course” is a well understood phrase in the field and does not need to be further defined. We also used hyperalignment to decompose the heritability of these traits into topography-dependent and independent portions, and (new to this version) also explicitly quantify the heritability of neural timescales, which we defined as the AUC of the ACF until the first negative ACF value in both the relevant Results and Methods sections.

      To make this clearer, we have modified the last paragraph of our Introduction to begin with:

      In the present work, we address these questions by analyzing 7T fMRI recordings of a twin sample acquired by the Human Connectome Project (Van Essen et al., 2013) to quantify the heritability of two distinct high-dimensional traits—stimulus-evoked BOLD time courses and functional connectivity profiles—across the cortex.

      Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      We believe that this question is getting at the difference between pairwise ISC (i.e., correlating one BOLD time course from one subject with that from another subject) and leave-one-subject-out ISC (i.e., correlating one BOLD time course from one subject with the corresponding average time course across all other subjects). We chose to use the pairwise ISC method because it allows us to capitalize on the information contained in the n<sup>2</sup> pairwise ISC matrix (whereas the other approach averages out meaningful information to yield a n<sup>1</sup> ISC matrix) and leverage a more sophisticated multidimensional heritability approach. Also, the leave-one-subject-out approach introduces additional issues re: handling family-level data (e.g., should we include a subject’s twin in the leave-one-subject-out average? If so, how should we handle subjects who don’t have a twin in the dataset, as averaging data from different numbers of subjects will lead to different ISC magnitudes? etc.).

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      We report p-values for heritability throughout our paper (e.g., stating that BOLD time courses are significantly heritable in 99% of parcels in Figure 2), and we believe that the reliability of our spatial maps across days of data collection (also quantified with p-values) further demonstrates the trustworthiness of our results. Finally, as we demonstrate in Figure S5, our sample size is more than sufficient to reliably detect small effects.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      In addition to driving shared neuronal responses (which are captured in BOLD signal oscillations <.1 Hz or so), movies also elicit shared cardiac, respiratory, and motion responses across participants at higher frequencies. Although we used a relatively conservative denoising approach here, we believe some of these non-neuronal signals are still present in our data; alternatively, it is also possible that these signals reflect “fast” BOLD responses at >.15 Hz (as discussed in 10.1016/j.neuroimage.2021.118658). In any case, the fact that information in this frequency band is considerably less heritable than information in slower frequency bands supports the idea that this band is noisier and suggests that our heritability results are driven by canonical neuronal activity-related BOLD signals.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      Although the decrease of 0.025 is small, we note that this constitutes around ~50% of BOLD time course heritability in some voxels (seen in comparison to Fig. 4C), and the spatial pattern of this result is quite consistent across days of data collection, indicating its reliability. Furthermore, the whole-brain distributions of results shown in Fig. 5B are clearly skewed towards negative values, indicating that controlling for NT partially reduces (or “explains”) BOLD time course heritability. Still, we agree that showing raw h<sup>2</sup> values in addition to the difference maps would be helpful for some readers and have added a corresponding supplementary figure (S12) which shows these.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      We did consider adding standard errors for these heritability estimates, but found that visualizing standard errors for each of the 153 unique network combinations in our heatmaps rendered the visualizations difficult to parse, and given that our hypotheses concerned global (e.g., hyperaligned vs. MSM-aligned) or network-level (e.g., sensory vs. associative) patterns, we focused on calculating standard errors/p-values for these analyses (although we note that dyad-level standard errors can be found in Fig. 6B, where they are clearly marginal compared to the group effects).

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      We agree that this result was relatively under-explored in our Discussion section and have added additional text (lines 851-855) to connect this result to recent work on arousal-dependent uniqueness of FC.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Do the authors have any ideas why we see this hotspot of heritability in pMTG/LOTC? It really jumps out in Figure 1A and Figure 2. The more posterior sensory MT+ area seems to drop when regressing out ISC in Figure 2B, but this pMTG area stays hot. Is there anything special about this kind of multimodal biological motion/action observation / social perception area (Pitcher & Ungerleider, 2021)? I don't think this is necessary to discuss in the manuscript, but I'm curious if the authors have any speculation.

      We are not certain as to why BOLD time courses in this parcel are particularly heritable- although this area is associated with biological motion, that particular function tends to be more right lateralized, and here we see nominally higher heritability in the left hemisphere. Per a Neurosynth review (and consistent with the left lateralization), we believe this may have more to do with speech processing, but a more definitive answer will require further investigation.

      (2) Page 3, line 127: "More information on these clips"-it might be worth saying a little bit more here just to make sure people understand that these are audiovisual clips, they include language, they're long enough to convey meaningful social and narrative information, etc.

      We agree and have added additional details on the clip composition to the relevant methods paragraph.

      (3) Figure 1 caption: can you add a sentence reminding readers what's going on with Day 1 and Day 2?

      We thank R1 for this suggestion and have added a sentence to this effect at this location.

      (4) Page 9, line 379: "although these more associative parcels do not encode a substantial amount of stimulus-specific information"-is this really true? I suspect these association areas still have decent ISCs, even if there are many processing stages downstream of the raw stimulus.

      Although these parcels are not the most synchronized by the stimulus, we agree that it is unfair (and vague) to say that they do not encode a substantial amount of stimulus-specific information. We have edited this sentence to make a more specific claim and highlight the relatively lower ISC in these parcels vs. more unimodal sensory areas.

      (5) Page 9, line 417: Can you unpack a bit more what you mean by "supra-BOLD frequency band"?

      Here, we refer to the fact that BOLD signals resulting from neuronal firing events have frequencies below ~.15 Hz (Josephs and Henson, 1999). We have added additional text and the Josephs and Henson citation to this line to further unpack this point.

      (6) Page 18, line 695: This discussion of how attention and gaze might partly shape response time series reminded me of recent work by Borovska & de Haas (2024)-might be worth citing.

      We are grateful to R1 for alerting us to this very relevant work and have included a reference to it in our discussion.

      (7) Page 19, line 755: I'm not sure I'd describe the hyperalignment results here as a "deleterious effects [on] heritability"-my reading was that hyperalignment allows you to say something more specific about heritability of function by allowing you to effectively factor out heritability effects that reduce to individual differences cortical topography; this seems like a good thing!

      We agree that “deleterious” was a poor word choice given its negative connotation, and have edited this sentence to read:

      “With this in mind, future studies investigating genetic correlations between brain function and behavioral variables may benefit from hyperalignment, as it can factor out individual-specific cortical topography and thus yield more precise estimates of functional heritability.”

      (8) I would love to see a ventral view in some of these plots! Not asking you to recreate the figures, but the ventral temporal cortex is an area of interest for many folks in the movie fMRI space (e.g., Haxby et al., 2011).

      We agree that ventral views would be of interest to some readers and have added the corresponding maps for our main results in supplementary figures S3 and S9.

      References:

      Borovska, P., & de Haas, B. (2024). Individual gaze shapes diverging neural representations. Proceedings of the National Academy of Sciences, 121(36), e2405602121. https://doi.org/10.1073/pnas.2405602121

      Haxby, J. V., Guntupalli, J. S., Connolly, A. C., Halchenko, Y. O., Conroy, B. R., Gobbini, M. I., Hanke, M., & Ramadge, P. J. (2011). A common, high-dimensional model of the representational space in human ventral temporal cortex. Neuron, 72(2), 404416. https://doi.org/10.1016/j.neuron.2011.08.026

      Pitcher, D., & Ungerleider, L. G. (2021). Evidence for a third visual pathway specialized for social perception. Trends in Cognitive Sciences, 25(2), 100-110. https://doi.org/10.1016/j.tics.2020.11.006

      Reviewer #2 (Recommendations for the authors):

      (1) To address the common core analytical problems listed under A), B), C), D), E), and basically throughout the methods:

      (a) Conduct permutations with exchangability restrictions to account for the pattern of dyad-relationships as e.g. implemented in PALM.

      (b) Control for age and sex covariates as covariates (e.g. as in SOLAR), rather than by matching.

      (c) Perform r-to-z transforms when conducting further analyses on correlations that assume normality.

      (d) For all analyses that assume normal distributions, e.g. in SOLAR and BrainSMASH, check that this is the case.

      We have explained how PALM is not suited for the study of effects that are defined at the dyad level (A), that we controlled for age and sex covariates in all our formal heritability analyses in our original submission (B), that we always performed r-to-z transforms when indicated in our original submission (C), and that our spatial permutation results don’t hinge on distributional differences (D).

      (2) Replace SEs derived from kacknife approach with those from SOLAR, or provide a comparison and motivation and/or demonstrate that SEs are correct.

      A more thorough explanation of the block jackknife procedure can be found in prior work introducing the multidimensional heritability method used here (Anderson et al., 2021).

      (3) Given problem (F & G):

      (a) Consider studying the parameters that drive the hyperalignment. They can be included as covariates in heritability analyses, and/or their heritability is of interest to understand the reasons for the heritability reduction post-hyperaligment.

      We agree that this would be interesting but the specific parameters that drive hyperalignment are beyond the scope of this study.

      (b) Include the alternative explanation of hyperalignment-induced noise in the discussion.

      We have added a figure showing that hyperalignment does not increase noise in ISC and explained here why “hyperalignment-induced noise” does not constitute a reasonable alternative explanation for our results.

      (4) Add heritability results for NT phenotypes.

      We have added heritability analyses for NT topography and (global) NT magnitude, as detailed above.

      (5) Motivate global signal removal, and acknowledge this process typically alters results substantially.

      We have added an explanation of our rationale for using GSR and shown in this response that it does not in fact substantially alter the results.

      (6) Rephrase and/or clarify the following:

      (a) "permutations quantify average differences" (under A).

      (b) "network combinations" and related analyses (under B & C).

      (c) why some analyses are separated per visit/day and others not (C).

      (d) methods and reasons for sample size estimation (C).

      We have rephrased or clarified all of the above.

      Reviewer #3 (Recommendations for the authors):

      (1) Participants should be recleared. I know HCP 7T data has 184 subjects. How can the authors have 176 twins and 690 unrelated subjects?

      As we reported in our Methods section, 178 subjects had complete movie-watching datasets, and 176 subjects had complete movie-watching and resting-state datasets. Of the 178 subjects with complete movie-watching data, we identified 690 age- and sex-matched dyads.

      (2) Figure 1. I don't find Figure S1A in Figure S1.

      We thank R3 for catching this error- we have amended this reference to read Fig. S1.

      (3) I could also suggest putting Figure 1 and Figure 2 together.

      We thank R3 for this suggestion- ultimately, we prefer to keep these figures separate to reinforce the difference between our dyadic similarity and formal heritability analyses.

    1. eLife Assessment

      This study presents important findings by identifying small molecules that can stabilize and refold missense-mutated VHL tumor suppressor protein, offering a potential therapeutic approach for clear cell renal cell carcinoma. The computational design approach is well-executed, but the evidence is incomplete due to insufficient demonstration that HIF2 downregulation occurs through on-target VHL rescue rather than off-target effects. Additional experiments with appropriate controls are needed to establish the specificity of the mechanism.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed some of comments raised in the previous round of review and have opted to proceed to a Version of Record without additional review.]

      Summary:

      This is an excellent and strong paper. The authors not only show the mechanisms of action of destabilizing mutations in VHL, but notably, they also go on to computationally design and experimentally test an inhibitor that restores wild-type pVHL function, offering starting points for a new class of kidney cancer drugs. The approach that the authors take here can be used to target destabilizing mutations in repressor proteins, common in diseases, including cancer.

      Strengths:

      This paper is the culmination of an extraordinary amount of work, over years, including method development and testing by a broad range of tools and experiments. It is thorough and comprehensive. It is also well-written and easy to follow.

    3. Reviewer #2 (Public review):

      Summary:

      Inactivating VHL mutations are common in clear cell renal cell carcinoma, and about half of those mutations unfold/destabilize the protein rather than directly interfering with critical protein-protein interactions. The authors identify a compound that can stabilize/refold mutant VHL and seemingly restore its ability to downregulate its major downstream targets.

      Strengths:

      The authors use a clever combination of virtual and cell-based screens, followed by suitable biophysical and cell-based validation assays, to arrive at a VHL refolder. This compound is suboptimal from an ADME point of view, but could be a starting point for further medicinal chemistry optimization. Success would have implications for other diseases linked to similar loss-of-function mutations.

      Weaknesses:

      In going from CP4 to CP4.29 the authors screened based on downregulation of HIF. This is logical but also introduces the danger of identifying chemicals that can downregulate HIF in an "off-target" manner i.e. non-specifically. It therefore essential to clearly show that CP4.29 downregulates steady-state levels of HIF and HIF target genes in cells with suitable (hydrophobic core) VHL mutants but not in isogenic cells lacking VHL.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We are most grateful to both reviewers for providing valuable feedback on our manuscript.

      Reviewer 1 had solely favorable comments, with no suggestions for revision.

      Reviewer 2 pointed out that experiment evaluating the effect of CP4 on pVHL half-life (originally included as Figure 3c) was difficult to evaluate because of CP4’s effect on pVHL abundance prior to cycloheximide treatment. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      Reviewer 2 also pointed out that experiment evaluating the effect of CP4.29 on HIF-2α half-life (originally included as Figure 4g) was not very compelling. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      We agree with Reviewer 2’s suggestion that additional experiments could further solidify that C4.29 downregulates HIF2 in a purely “on-target” manner, however we prefer to reserve such studies for the future.

      Reviewer 2 also made several valuable suggestions for the text itself (awkward wordings / citations / clearer figure legends). We appreciate this feedback and have updated the text accordingly.

    1. eLife Assessment

      This important study advances our understanding of the biomechanics of seed processing in birds by providing a comprehensive 3D kinematic analysis of coordinated bill and tongue movements across two species with contrasting biting forces. The evidence is convincing, combining high-speed XROMM with Bayesian statistical modeling in a rigorous and technically innovative framework that advances the understanding of avian feeding kinematics. Strengthening the statistical validation of qualitative claims, particularly for tongue-seed velocity relationships, and improving the accessibility of the probabilistic modeling framework would further solidify the conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      The authors quantified and compared the 3D kinematics of bill and tongue movements between two seed-eating bird species: one that specializes on soft seeds, and one that is more adapted to feeding on hard seeds. Their goal was to determine specifically what the role of the tongue was for processing (e.g., dehusking) seeds, and to understand how differences in biting strength between species affect other aspects of seed processing. The authors provided intricate (visual) details of seed processing movements, and showed how coordination between the tongue and cranial kinesis (i.e., mobility of the upper bill relative to the cranium) is both critically important for properly positioning seeds to enhance feeding efficiency. Many studies have detailed how seed-eating birds process seeds, but this study has elevated those to a new level of quantification and visualization for readers to fully experience firsthand. Furthermore, the authors established that the force-velocity trade-off that has been observed between bill functions (e.g., feeding and singing) is largely driven by the contractile properties of the muscles. The conclusions are well supported by the results, and the authors placed the results more broadly into the context of manual grasping, making the argument that these birds achieve high levels of dexterity with far fewer degrees of freedom, which could have potential biomimetic applications.

      Strengths:

      This study builds upon - and advances - our understanding of the feeding mechanics of seed-eating birds using cutting-edge 3-dimensional modeling and kinematics. Their quantitative analyses of upper and lower bill, tongue, and seed displacements are complemented by elegant visualizations of seed processing in each species. Their comprehensive Bayesian modeling statistical framework tackles the issue of small sample sizes (i.e., few subjects) with volumes of data for each (i.e., lots of sequential kinematic variables) that plague comparative biomechanics studies, principally because (a) it is difficult to gather these high resolution XROMM and muscle contractile data on more than just a few subjects, and (b) these data streams are inherently very large, as they are gathered at high frame and sampling rates. Furthermore, I believe their approach to statistically testing for differences between species sets a new standard for our field that could (perhaps should?) be implemented in other similar types of studies. Another strength is in how the results were packaged: each subsection indicated how the objectives were addressed, and there were concluding statements trailing each subsection that helped deliver the key takeaways.

      Weaknesses:

      A potential weakness is one that the authors themselves mentioned, regarding the body (and skull) size differences between species. Because gape size limits bite force, and given the force-velocity tradeoff in muscle function, there could be limitations on the rapid manipulation of relatively large seeds for similar reasons in the smaller finches. I see that the small finches appear to overcompensate in their beak rotations, but it's not clear how those compensatory movements might affect their seed processing kinematics with their preferred seed sizes. This does not nullify the authors' conclusions, but the results for the smaller finches might not be entirely representative of seed processing mechanics in smaller species.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates coordinated beak-tongue movements in seed manipulation, biting, and dehusking in songbirds. A comparative analysis of the seed-eating process in two songbird species with different biting forces, the domestic canary and Java sparrow, was conducted using high-speed XROMM with anatomical marker tracking and quantitative behavioral analysis. The authors have done a great job analyzing upper and lower beak rotation and translation, seed orientation and movement speed, and tongue kinematics.

      Strengths:

      The methodological approach of using high-speed (500 fps) X-ray reconstruction for 3D kinematic tracking in small animals is novel and powerful. It enables high temporal resolution tracking of orofacial movements and could potentially inspire future orofacial research in mammals, including mice and marmosets. Moreover, this study encompasses a wide range of anatomical components involved in seed manipulation behavior, including the upper and lower beak, the tongue, and jaw muscles. The behavioral quantification of these components is solid. The findings that both the upper and lower beaks contribute to seed processing, that the lower beak exhibits greater up-and-down and left-to-right flexibility than the upper beak during seed processing, and that the tongue plays an important role in transporting seeds into the mouth are all solid conclusions consistent with observations of bird feeding behavior. Nevertheless, it is valuable to confirm and quantitatively characterize these observations experimentally. The videos are excellent and very informative.

      Weaknesses:

      (1) The paper often resorts to qualitative descriptions (e.g., "a high positive correlation of tongue velocity and seed velocity", "Compared to positioning, the measured velocities of both seed and tongue were much lower") instead of providing exact quantitative measurements or statistical results. The authors stated that temporal autocorrelation biases standard statistical analyses (lines 205-210), but this rationale does not justify the absence of statistical validation. Suggestion: use appropriate methods for time-series data, such as a permutation test, to test the significance of correlations between variables and avoid false positives.

      (2) (Minor) The marker-tracking image shown in Figure 1B could benefit from the inclusion of a higher-contrast, zoomed-in frame of the head showing the metal markers without the red tracking points, alongside the same frame with the red tracking points overlaid, to provide readers with a clearer view of the X-ray image and the methodology and its precision.

      (3) (Minor: possibly soften the mechanistic claim). The proposed mechanism of lingual papillae on the tongue surface may aid food manipulation and food movement towards the posterior region of the mouth is interesting, yet the evidence describing their morphology is not strong enough to support the claim about their functional roles. Furthermore, the claim that papillae orientation affects food transport in lines 294-296 lacks supporting experimental evidence. In addition, the roles of extrinsic and intrinsic tongue muscles in controlling dexterous tongue shape changes and movements are not discussed.

    4. Author response:

      We would like to express our gratitude for the thorough evaluation of our manuscript by the editors and reviewers. We are grateful for the overall positive assessment. The suggestions for improvement are reasonable, and we are certain that addressing these points will improve the clarity, accessibility, and scientific integrity of the study. Thus, we plan to conduct a revision of the manuscript, addressing all the points raised. The most important planned adjustments are outlined below.

      (1) Improving the accessibility of the probabilistic modeling framework

      Reviewer 1 kindly stated that our Bayesian modeling framework for testing for species differences 'sets a new standard for our field.' As a new standard, however, the method should be explained in a more accessible way. Hence, we plan to provide additional explanations for the statistical workflow, e.g., by providing comprehensible visuals, to make the workflow easier to understand and easier to apply.

      (2) Statistical validation of qualitative claims

      We acknowledge that a statistical validation of qualitative claims regarding the relationship between seed and tongue movements and between upper and lower beak movements would considerably strengthen the validity of our findings. We thank Reviewer 2 for bringing permutation tests to our attention for quantifying the correlation between time series. Since permutation tests involving index-shuffling of one of the data sets are generally not valid for time-series data [1, 2], we'll consider a variant of a trial-swapping permutation test, such as a permute-match test [3]. Alternatively, the truncated time shift (TTS) test [2] might be an option, as also this method is valid for auto-correlated time series data. At this point, we can't tell yet which method we'll use for the revised manuscript. We need more time to assess the requirements of each method and evaluate which test is most appropriate to answer our specific research questions and best fits our kind of data.

      (3) Adjustments in the discussion

      Following the suggestion by Reviewer 1, we'll refine our discussion on the effects of skull size differences, putting more emphasis on the implications of potential effects for feeding kinematics in small species.

      Furthermore, as suggested by Reviewer 2, we'll soften our discussion on potential functions of lingual papillae in seed processing, as the current literature lacks experimental evidence for the claimed mechanistic roles.

      References

      (1) Yuan, A. E., & Shou, W. (2022). Data-driven causal analysis of observational biological time series. Elife, 11, e72518.

      (2) Yuan, A. E., & Shou, W. (2024). A rigorous and versatile statistical test for correlations between stationary time series. PLoS biology, 22(8), e3002758.

      (3) Yuan, A. E., & Shou, W. (2025). Permute-match tests: Detecting significant correlations between time series despite nonstationarity and limited replicates. eLife, 14.

    1. eLife Assessment

      This valuable study investigates the neural basis for recovery of complex wheel running behaviour following a unilateral spinal cord injury in mice. By combining behavioural analyses, whole-brain mapping, and tracing techniques, the authors provide incomplete evidence that new cortico-medullary connections can drive effective motor recovery. The paper could be strengthened with manipulations to establish causality, a more fine-grained analysis of the behaviour, and some reorganisation of how the data are presented and discussed.

    2. Reviewer #1 (Public review):

      Summary:

      The authors seek to understand and identify the neural plasticity that underlies recovery from precise unilateral hemi-pyramidotomy. The corticospinal tract is severed on one side in the pyramids below the exit of corticoreticular projections. Recovery from the injury is achieved with an intensive wheel running rehabilitation regime. The anatomical sites of plasticity, the importance of plasticity in different reticular areas<br /> to recovery, and the impact of the degree of plasticity observed on recovery as correlated predictors, are shown.

      Strengths:

      Refined anatomical analysis using mouse line and genetic and viral intersectional tracing identifies specific reticular targets of likely enhanced cortical control that correlate with recovery of locomotor skill.

      Weaknesses:

      (1) The study is correlational at this time. This does not undercut the value of the data and the identification of targets of plasticity achieved in the work.

      (2) Generalization of motor gains beyond locomotion was not tested. Reach-to-grasp tasks for feeding were not tested.

      (3) Some discussions and use of the terms fine motor and skilled motor are fuzzy, and the limitations of the study are not sufficiently clearly stated.

    3. Reviewer #2 (Public review):

      Summary:

      Bonanno and colleagues combine unilateral pyramidotomy, continuous voluntary complex-wheel running, whole-brain intersectional CSN tracing, and c-Fos mapping to ask whether rehabilitation reorganizes the supraspinal collaterals of the intact corticospinal tract neurons. The study is technically ambitious and competent, the uPyX + complex-wheel + intersectional-tracing + BrainJ combination is smart and interesting, the behavioral effect is convincing, and the blinding and exclusion criteria are explicit. The central anatomical finding - a CSN-specific, whole-brain projectome comparison with subregional LPGi/GiA/MdV granularity - is a legitimate contribution that builds on Asboth 2018. However, the strength of evidence does not support the strongest causal wording in the current abstract, significance statement, and parts of the discussion: the results remain correlational, the MdV-behavior correlation is modest, and its significance is sensitive to the unit of analysis. A major revision is recommended, primarily of framing and quantitative robustness, rather than because the central dataset is unconvincing.

      Strengths:

      (1) Technically ambitious and technically competent study addressing a relevant gap: brain-wide mapping of intact-CSN reorganization under continuous voluntary rehabilitation.

      (2) The combination of uPyX, complex-wheel running, intersectional tracing, and BrainJ whole-brain projection analysis is novel and well integrated.

      (3) Behavioral effect is convincing, blinding, and exclusion criteria are explicit.

      (4) The central anatomical finding (CSN-specific whole-brain projectome under rehab, with LPGi/GiA/MdV subregional resolution) is a legitimate contribution that builds on Asboth 2018. The closest recent works (Lemieux et al. 2024, Jeleva et al. 2026) study reticulospinal rather than CSN plasticity and are complementary rather than competing.

      Weaknesses:

      (1) Causal framing extends beyond what the current evidence supports.

      The abstract and significance statement present MdV as a potential mediator, or even a central locus, through which rehabilitation re-establishes descending control of the impaired limb. This is stronger than the evidence. What the paper shows is that CSN collateral projection density in MdV has a mild-to-medium correlation with behavioral recovery, and that this region is already known from prior work (Esposito 2014) to be relevant for skilled forelimb function. That is an interesting anatomical correlation, not a demonstration of mediation. No manipulation of MdV or of MdV-projecting CST terminals is performed; there is no silencing, no pathway-specific perturbation during rehabilitation, and no test showing that the identified sprouting is necessary for recovery. The limitations section acknowledges this, but the prominent claims do not.

      (2) The behavioral caveat on what is actually novel.

      The cleanest way to state what is genuinely new, clearer than the abstract itself, is this: when a CSN population loses part of its spinal target domain (via contralateral uPyX denervating the opposite cord), some CSNs from the opposite cortex appear to redirect growth into brainstem collaterals (LPGi, GiA, MdV). The compensation is plausibly sufficient to restore gross descending drive to the impaired forelimb, but most probably inadequate for the fractionated, cortico-motoneuronal fine-grain control that the direct CST normally provides. That distinction - recovery of drive and even skilled locomotor control vs. recovery of fine precision - is consistent with the ladder-rung improvements the paper reports (footfall counts are an integrated gross-placement metric) and with the skilled-reaching literature (Esposito 2014 and similar), which suggests precision grip and digit individuation would not be fully recovered by an MdV-centered detour. This note is also translationally important when we ask what humans consider fine motor control, which is mostly object manipulation. Relatedly, the ladder task is "skilled" in the operational sense that it requires cortical control, but the motor output measured (gross paw placement, overreach) is not fine motor function in the sense of digit individuation, grip force modulation, or pellet manipulation. "Skilled" here does not even mean *acquired* skill: classical skilled reaching in rodents involves explicit training to acquire a novel motor program, whereas here mice are only habituated. The brainstem-compensation hypothesis is more comfortable with restoring cortex-dependent gross placement than with restoring acquired fine-motor skills.

      (3) The anatomy sample is modest for the precision of the claims.

      Projection analysis rests on n = 9 pooled controls, n = 5 uPyX−Rehab, and n = 5 uPyX+Rehab. For a whole-brain subregion analysis, this is not a large dataset, even with the sensible restriction to the Wang et al. spinally-projecting set. The three medullary hits are plausible, but some of the most specific conclusions rely on a relatively small number of animals for its most specific claims. This matters especially for the MdV-behavior correlation.

      (4) Normalization enforces a zero-sum structure.

      Projection density is normalized to the total CST tract signal. This is a reasonable way to control for tracing variability, but it imposes a relative structure on the data: an apparent increase in one region may partly force an apparent decrease elsewhere. This may matter and has to be looked into by the authors, because the manuscript interprets decreased density in some other targets as meaningful redistribution.

      (5) The decision to merge PMn and MdV under a single "MdV" label needs more justification.

      Since the discussion relies on prior literature assigning skilled forelimb function to MdV proper, the reader needs to know whether the signal truly localizes there or whether it may partly reflect a neighboring region grouped under the same atlas label. Related to this, laterality would be very informative: since the proposed compensatory route is anatomically directional, showing whether the increased signal is preferentially located on the expected side of the medulla would strengthen the interpretation.

      (6) The c-Fos / Fig. 3 section goes beyond what the data directly support.

      The section "Complex-wheel running recruits intact corticospinal neurons" and the figure title "Rehabilitation functionally recruits intact CSNs" go beyond the actual observation, which is that a higher fraction of CSNs in M1 and M2 are c-Fos+ in runners than in non-runners. "Functionally" is not supported: c-Fos is a transcriptional marker of recent activity, not a functional readout; it does not show that the CSN's output is used to drive behavior. "Rehabilitation" is not supported either: the contrast is runners vs non-runners, applied uniformly across Sham and uPyX groups - healthy Sham+Rehab animals are on wheels for leisure, and the c-Fos effect is present in them too. The finding is difficult to interpret without thinking of the simpler framing ("moving mice have more motor cortex activity than resting mice"), with no control for generic arousal or ambulation. This section is the softest link in the causal chain running - CSN activity - medullary sprouting - recovery.

      (7) MdV-recovery correlation: unstated multiple-comparison correction and possible pseudoreplication.

      The correlation (R² ≈ 0.33, p ≈ 0.01) is the backbone of the paper's "causal" claim. Panels L/M/N test three correlations (LPGi, GiA, MdV vs forelimb footfall recovery); only MdV is reported as significant. The Figure 5 legend applies Tukey adjustment to the t-tests in A-C but makes no analogous statement for the correlations in L-N. A 3-test Bonferroni (α = 0.017) would not flip the MdV result, but disclosure is warranted, and the three tested regions were pre-selected from the significant group contrasts in A-C, which, to a statistician, would further shrink effective α. More importantly, the figure legend states that closed and open circles represent CFA- and RFA-traced values, respectively, which suggests the correlation treats the two tracer channels per mouse as independent datapoints - doubling the apparent n (≈ 20 from 10 uPyX mice), with the result of a higher significance than one would have at the mouse level.

      (8) Reporting issues.

      The reader would benefit from judging statistical choices such as those above directly from a data table rather than interpreting the authors' choices. The SciScore rightfully flags multiple missing components of transparent reporting: missing RRIDs, no code availability, limited data availability, and no power calculation, among others.

      Almost all these weaknesses can be addressed with a revision of the manuscript, especially in the framing of results.

      Conclusion:

      The core message - that rehabilitation is associated with a selective pattern of CSN collateral remodeling in the motor medulla, and that MdV projection density covaries with behavioral recovery - is defensible from the data and already a useful result. The current wording in parts of the abstract, significance statement, and discussion goes beyond this and implies a mechanistic conclusion (mediation, central locus, re-establishment of descending control) that the data do not yet establish. The manuscript would better match its evidence with "associated with", "correlates with", or "candidate locus" framing, unless a causal experiment is added.

    4. Reviewer #3 (Public review):

      Summary:

      In this study, Bonanno et al. show that after a lesion of the corticospinal tract (CST), rehabilitation running in a complex wheel drives improvement in skilled forelimb performance in mice. Mice with unilateral CST injury can perform gross motor tasks (locomotion) at the same level as the non-injured mice, but injured mice still have deficits in another task involving fine motor control. Thus, it is well-suited to test the efficacy of locomotion-based rehabilitation in fine motor control. Mice that voluntarily engaged in the rehabilitation protocol improved in the fine motor control task more than those mice that did not perform any rehabilitation. Highlighting the role of rehabilitation in the recovery of motor function after the lesion.

      The authors aimed to study rehabilitation-driven intact CST sprouting to supraspinal areas. They identified one area in the motor medulla where rehabilitation significantly changes the projection density from the intact cortical spinal neurons. Interestingly, this area has ipsilateral connections and thus could be a pathway to convey motor commands from the intact corticospinal tract to the denervated area. However, as the authors acknowledge in the discussion, they only found a correlation between the change in the synaptic projections from intact CST to the medulla and the recovery. Future work should study if indeed the area of the motor medulla identified here increases its ipsilateral projections to the denervated area, confirming the re-routing of motor commands from the intact cortico spinal tract to the denervated area. The paper is strong and, in general, claims are supported by the data.

      Strengths:

      In this study, Bonanno et al. show that after a unilateral corticospinal tract lesion (CST), locomotion rehabilitation can improve motor function and improvements generalized to tasks that require fine motor control. Moreover, it identifies a potential pathway that could be used for the intact corticospinal tract to convey motor commands to the denervated area. The pathway identified here could become a target for rehabilitation therapies.

      Weaknesses:

      As the authors acknowledge in the discussion of the study, the main limitation of this study is that the reorganization observed at the motor medulla is only correlational. Thus, it is possible that the adaptation to running with an injured limb of the intact CST to adapt to an injured limb rather than a re-routing of the intact CST inputs to the denervated area underlies the synaptic changes observed in the motor medulla.

      The statistical analysis could be better described.

      The generalization of skilled movement is limited to only locomotion tasks.

    1. eLife Assessment

      The worldwide decline in the health of coral reefs is well documented, and overgrowth by microbial consortia can be a contributing factor. Kelman and colleagues used metagenomic analysis to interrogate potential changes in phage-associated genes predicted to be involved in central carbon metabolism. The study addresses the hypothesis that metabolic genes associated with carbon metabolism that are encoded by viruses reflect the health of the corals. The study contributes a valuable perspective on the potential role of phages in coral health, although limitations of the data and analyses offer an exploratory examination rather than a definitive result. Overall, the evidence supporting the major findings is incomplete, in part because the conceptual model relies on qualitative assumptions rather than empirical data.

    2. Reviewer #1 (Public review):

      Summary:

      Microbialization (bacterial overgrowth) is a recognized component of degraded, eutrophied coral reefs where there is a shift from coral to algal dominance on the benthos. In addition, previous work has demonstrated that virus communities shift from a lytic strategy dominated (kill-the-winner) to a temperate (lysogenic) strategy dominated with reef microbialization. Kelman et al. sought to leverage previously published virus metagenomes produced from the water column of healthy and degraded coral reefs to assess virus community metabolic shifts. The authors also produce a conceptual model to demonstrate the potential impact of the observed metabolism shifts on reef fates.

      Strengths:

      The main strength of the manuscript is the findings from their metagenomic analyses and results. The virus metagenomes were produced using established approaches in the field and yield sufficient data per sample for their analyses. Interesting results regarding the shift in the types of genes from anaplerotic to cataplerotic provide the foundation for testable hypotheses to determine the magnitude of impact virus strategies have on reef health. The introduction is also well written and sets up the scene very well.

      Weaknesses:

      (1) The methods text currently omits important information related to the sampling design. It is not clear how many metagenomes are from healthy and degraded communities. This impacts the interpretability and robustness of the statistical results. Furthermore, it is unclear if analyses are based on assembled contigs or read-based alignments. Improving the clarity and organization of the Methods is essential for reproducibility.

      (2) Regarding the bioinformatics approach, normalization using the "percent known" approach within samples may not fully account for discovery bias related to sequencing depth. While Supplementary Table 1 shows variability in read counts, the lack of community-level metadata makes it difficult to determine if sequencing depth covaries with community type (healthy vs. degraded). The study would benefit from a rarefaction analysis or subsampling to ensure that gene frequency trends and Spearman correlations are biological signals rather than artifacts of sequencing effort.

      (3) The qualitative model in Figure 5 is positioned as evidence for the role of viruses in reef health, but it does not provide independent support for the authors' hypotheses. Since the model is parameterized using "arbitrary units" to reflect the authors' assumptions rather than being derived from the empirical metagenomic data, it serves as a helpful illustration of a hypothesis but not as a validation of the findings.

      (4) Results and discussion require revisions to improve readability and connectivity across sections. Ensuring a clear distinction between empirical data and model-based speculation would help the audience better appreciate the science.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Kelman and coauthors investigates how viral communities differ in the genes they encode in healthy and degraded coral reef ecosystems. Across 19 viral metagenomes from Central Pacific reefs, the authors assess the frequency of integration/excision genes as a proxy for viral community temperateness and ask whether genes associated with central carbon metabolism covary with signatures of temperateness. The main finding is that viral communities with more temperate-related genes encode more genes from the Entner-Doudoroff pathway and other reactions interpreted as anaplerotic, whereas more lytic-associated viral communities show greater representation of some pentose phosphate pathway, TCA, and redox-associated genes interpreted as cataplerotic. The authors propose a model based on these patterns in which lytic viral metabolism helps suppress bacterial overgrowth on healthy reefs, while temperate viral metabolism may promote microbialization on degraded reefs. The study addresses an interesting and potentially important concept - that viral auxiliary metabolic genes are important components of microbial communities and can affect ecosystem functioning. Linking viral metabolism to coral reef microbialization is a creative conceptual advance. The manuscript is clearly written, and the reported enrichment of anaplerotic genes in temperate-associated viromes is an interesting pattern that could motivate future work on how viral metabolic potential varies across reef states.

      Strengths:

      (1) The study connects viral lifestyle, central carbon metabolism, bacterial overgrowth, and reef degradation in a framework that could be useful for future studies of coral reef ecosystems and viral ecology. This is an interesting synthesis that links viral auxiliary metabolism to broader questions about microbialization and reef state.

      (2) The manuscript is generally clearly organized around a testable prediction: viral metabolic gene content should vary along a lytic-to-temperate viral community gradient. The reported enrichment of anaplerotic genes in viromes with a larger fraction of temperate viruses is a compelling result.

      (3) The authors highlight several virus-encoded metabolic genes that may not have been previously reported in viral datasets or genomes. If supported by further validation, these observations could expand the known repertoire of viral metabolic potential.

      (4) The modeling helps clarify the feedbacks the authors propose may connect viral lifestyle, bacterial metabolism, and coral reef degradation. It provides a foundation for generating hypotheses about how viral metabolic genes could influence reef microbial dynamics.

      Weaknesses:

      (1) The main limitation is that the evidence for several key claims remains indirect. The core analysis is based on correlations between metabolic gene frequencies and integration/excision-related genes. This does not demonstrate that the metabolic genes occur in temperate viral genomes, are physically linked to lysogeny genes, are expressed during infection, or alter host metabolism. Thus, the data support an association between VLP-associated metabolic annotations and a community-level temperateness proxy, but not a direct link between temperate phages and these metabolic functions.

      (2) It is important not to equate community-level gene frequencies with genome-level or infection-level metabolic programs. A virome may contain more anaplerotic genes overall, but that does not demonstrate that individual viruses reprogram their hosts in an anaplerotic manner nor that infection produces a net anaplerotic effect. Individual viruses may encode both anaplerotic and cataplerotic genes, and a smaller number of cataplerotic genes could have stronger metabolic consequences depending on expression, enzyme efficiency, pathway position, and host context. This is an important limitation that should be acknowledged and, if possible, addressed with contig- or genome-level analyses.

      (3) The ecological interpretation assumes that viral infection is strong enough to influence reef-scale bacterial population dynamics. However, the study does not directly measure infection frequency, lysis rates, viral production, burst size, lysogeny frequency, prophage induction, gene expression, or bacterial mortality. If viral mortality or lysogenic conversion were rare in these systems, the observed gene-frequency patterns could have limited ecosystem-level consequences. This makes claims about viral metabolism suppressing bacterial overgrowth, accelerating microbialization, or acting as a conservation lever more speculative than suggested.

      (4) There are statistical limitations related to the use of relative gene frequencies. Because genes are normalized as percentages of known genes, the data are compositional. Apparent increases in some categories may partly reflect decreases in others. Bootstrapped Spearman correlations are useful for assessing the robustness of these associations, but they do not address compositionality or multiple testing.

      (5) The anaplerotic/cataplerotic classification is central to the manuscript's conclusions and would benefit from more support. The framework is useful, but it depends on both annotation confidence and biochemical context. Sequence-similarity annotations alone may be vulnerable to misannotation, especially for central metabolic enzymes that share conserved domains across functionally distinct proteins. Stronger evidence that key genes contain key functional domains and/or are phylogenetically related to characterized enzymes would help support the proposed functions. In addition, many central carbon enzymes are reversible or context-dependent, so a clearer rationale for each classification would strengthen the interpretation.

      Overall, the manuscript presents a valuable hypothesis and highlights new ecological patterns in coral reef viral metagenomes, but falls short of the evidence needed for the strongest claims. The work would be strengthened by analyses that directly link metabolic genes to viral genomes or lysogeny markers, address compositional effects, validate key annotations, and more clearly distinguish observed gene-frequency associations from hypothesized effects on infection, host metabolism, and reef state.

    1. eLife Assessment

      This Review Article puts forth a normative theory for the grid cell representations found in the entorhinal cortex. It discusses a range of theoretical models and experimental findings, organizing them around a proposed framework in which grid cells are interpreted as biologically constrained, high-fidelity codes for path integration. This framing can be potentially interesting both for readers seeking a conceptual entry point into the grid cell literature and for those more generally interested in the promises and limitations of normative theories in neuroscience. Some logical gaps and points requiring conceptual or technical clarification were nonetheless identified. Moreover, the empirical support for the path-integration account is not yet as definitive as the manuscript's framing sometimes suggests. The review would thus be strengthened by clearer justification of key arguments and fuller discussion of biological complexities, model limitations, and competing interpretations. Some stylistic choices in how arguments and literature are sometimes rhetorically framed may lessen the review's appeal for key segments of its intended audience.

    2. Reviewer #1 (Public review):

      Summary:

      The review by Dorrell and Whittington synthesizes the progress made over the past few years with respect to a normative theory of grid cells. The core question addressed by normative frameworks of grid cells is what primary computational function grid cells serve. The review discusses evidence from mechanistic models and experimental data that point to path integration as the computational function of grid cells, consistent with results from normative models. The main goal of the review is to clarify the normative grid cell theory literature. However, the current version of the article reads at times more like a perspective or opinion article in support of the path integration hypothesis rather than a critical review of normative frameworks in the grid cell literature that contrasts the benefits and limitations, as well as pitfalls and caveats, with other modelling approaches.

      Some specific comments are as follows:

      (1) Abstract: "The first question quickly attracted an answer: grid cells subserve path integration ..." - I am not sure if this statement is correct. The first grid cell paper by Hafting and Fyhn in 2005 suggested that grid cells are part of a path integration-based map, and the paper emphasizes the map part. It remained unclear, and is still debated, whether grid cells are part of a system performing path integration or whether grid maps reflect the output/result of a path integration process. Other theories about the function of grid cells were brought forward as well. Although the main competing theory is discussed in this review, this review article at times appears more as a perspective or opinion article with a clear bias toward the path integration hypothesis rather than objectively discussing the evidence.

      (2) Grid cells may serve multiple functions. What would be the implications for our understanding of grid cells and for interpreting the results of normative models? In general, the review could discuss some pitfalls or caveats of normative models in more detail.

      (3) A normative framework can be helpful in two ways: (a) Given sufficient details on biological constraints, a normative model can help identify the computational function of grid cells. If a computational function is given and - under the given simulated biological constraints - grid cells were part of the solution, the results of the model would support the hypothesis that grid cells serve the computational function in question. (b) If a computational function were identified beyond any doubt (e.g., assume experimental data demonstrated that grid cells are necessary and sufficient for path integration), a normative model would help identify biological parameters necessary to produce grid cell firing. Unfortunately, the review falls short in making this clear distinction between (a) and (b) and in discussing important caveats regarding mixing up these two ways. E.g., the neural network model approaches by Sorscher et al. and others have been criticized because they try to achieve two things at the same time: find support for the computational function of grid cells and identify optimal parameters that result in grid cells. But doing both at the same time provides a strong bias in tweaking the parameters in exactly the way you need for the model to produce grid cells as a solution (other solutions may be possible given other parameters), preventing strong conclusions regarding the computational function of grid cells and preventing conclusions about what the parameter choices mean for biological connectivity motifs. These caveats in setting up normative models and interpreting them could be discussed in greater detail.

      (4) A common assumption underlying most grid cell models is that head direction is viewed as identical to movement direction. However, head direction can differ at times from movement direction, and entorhinal head direction cells code head direction rather than movement direction (Raudies et al., 2015; 10.1016/j.brainres.2014.10.053). This missing link in how movement direction signals reach and inform grid cells could be discussed.

      (5) "Knowing that one neuron in a module is active and that you make a movement north uniquely determines which neuron in that module should be active next" - I agree that this rule follows from the fact that grid cells within one module differ in phase but share spacing and orientation. However, I am surprised that the authors do not also make the argument here for the value of a normative model. Rebecca R.G. et al. (10.7554/eLife.96627) use exactly the rule cited above as a normative function. They demonstrate that this rule begets grid cells. Isn't this a prime example of how a normative approach can contribute to scientific inquiry? First, a hypothesis about a computational function is derived from experimental data. And in turn, using a normative framework, the experimental data are derived from the computational function (under appropriate biological results). The paper is discussed later together with Nicolai Waniek's work (10.1162/neco_a_01255). However, in my opinion, their work seems to be somewhat misrepresented in that later paragraph. E.g., velocity is still required as an input to determine which neuron should be active next, neurons do not need to be binary units, and space is not discretized beyond the fact that space is encoded by neurons with spatial firing fields.

    3. Reviewer #2 (Public review):

      Summary:

      This review by Dorrell and Whittington covers a number of aspects related to normative modeling of grid cells. They begin by discussing key experimental insights on grid cell phenomenology. Then, they discuss how grid cells can be used to perform path integration and how they size up as efficient codes of space. These two sections then lead the authors to discuss how combining path integration and efficient coding objectives leads to models of axis-aligned grid cells in a single module. Discussion on non-linear objectives leading to multi-modules is presented. The review ends with several outstanding questions and an optimistic outlook of how normative models (particularly, task-optimized RNNs) can be used as tools for advancing understanding in neuroscience.

      Strengths:

      (1) The review is timely and covers an area that has seen a lot of recent activity. This discussion around many of the different results (and kinds of models), I think, will be generally helpful for the field.

      (2) Although I think the story could be a little more coherently made (see below), in general I enjoyed the author's flow from efficient coding -> efficient coding + path integration -> efficient coding + path integration + non-linear objective. This framing supports the specific conclusion the authors arrive at.

      (3) I also really liked the message that the review made of how normative modeling, despite some of its challenges/limitations, can be used effectively in neuroscience. The discussion of cycling between "experimental" modeling (e.g., vanilla RNNs) and theoretically-grounded models was nice, and I think it helps demonstrate the value of this approach.

      (4) Showing how the metric loss could be seen as a bandpass filter (Figure 3C) was nice and a contribution of the review.

      (5) While the focus of P4 (conjunctive HD-grid cells) felt initially a little cast aside, the discussion around "brain and task-optimised RNNs with standard architectural choices use fundamentally different path-integration mechanism" was nice and I think helpful for steering the community to an interesting open problem.

      (6) Identifying how "non-linear functionality" can lead to multi-modules was nice and not something that I have seen as clearly presented before.

      Weaknesses:

      (1) The authors view the experimental evidence for grid cells being linked to path integration as "specific and strong" and that the " key computational feature that defines entorhinal cortex [is] path-integration". I think experimentalists (at least the ones I work with) would push back on that. First, it's hard to isolate path integration in rodent experiments. So while Gil et al. (2018) did about as good a job as you could do, there are still other interpretations of the results that are not purely path integration dependent. And second, as the authors point out later in the review, there is experimental work finding that grid cells are disrupted in large environments and 3D. Path integration certainly happens (to some extent) in these spaces, which begs the question of how it is achieved with weakened grid coding. Thus, I think reducing the claims about how strongly grid cells are experimentally linked to path integration is called for.

      (2) The authors introduce the idea of efficient coding of space and discuss how grid cells are not optimal. It is later clarified (Sec. 5.3) that multi-module codes can be efficient (even if not the most optimal). I was confused reading Section 3, because in Section 2 the multiple modules are discussed, but then in Section 3, they are dropped, and only a single module is being considered. Equation 2 was also a little confusing to me. Alpha is not defined, and I would have thought that it would be x^Tx' - g(x)^T g(x') and not x^Tx' g(x)^T g(x'). Given that there is no page limit here, I think a little more detail in Section 3 would be helpful.

      (3) In Section 3, the authors make use of P2 (translation invariance within a module) to rule out (or, at least, question) certain models/approaches. While this is certainly a standard assumption made in theoretical work, it is not very well supported by experimental findings. In particular, Diehl et al. (2017), Ismakov et al. (2017), and Dunn et al. (2017) all found that individual grid fields systematically vary in their peak firing rate. In addition, Redman et al. (2025) found that, within a given module, there was a small but robust diversity of grid orientations and spacings. These suggest that grid cells within a single module may actually be able to encode properties of local space and give some support to normative models that find efficient space coding with grid cells by finding non-axis-aligned grid fields. I think this is all important to mention because: a) it provides more biological nuance to the question about spatial coding; b) it provides more ways in which to test models. For instance, in Redman et al. (2025), the Sorscher et al. (2022) model was shown to produce variability in grid properties that loosely matched what was found in real data. For tests like this (e.g., how much does a model reproduce variability in grid firing field peak rates), I think it is going to be important for continuing to evaluate models.

      (4) The focus of the review, I know, is grid cells, but of course, grid cells are part of the MEC and the larger hippocampal network. I totally understand, at some level, you have to make a decision of what to model, but it seems that there are other functional classes of neurons (border cells, head direction cells) that all play an important role in path integration. And while the models the authors consider at the end of the review capture properties of grid cells really well, they do so at the cost of not modeling anything else. The authors mention this in the context of the models not capturing conjunctive grid-head direction cells, but I think the point is a deeper one, and more discussion of at what level it makes sense to consider grid cells only is important.

      (5) As I mentioned in the Strengths section, I did enjoy the flow of the paper on how path integration + efficiency is needed to get grid single modules and path integration + efficiency + non-linearity is needed to get multiple grid modules. This creates the story that adding more of these theory-driven constraints helps lead to more "accurate" models of grid cells. But one alternative view is that, if path integration + efficiency is enough to get a single grid module (but only a single grid module), then maybe the utility (or need) of multiple grid modules comes from something else. That is, instead of saying "we need more constraints to get multiple modules", it could be evidence for "we need to re-think whether multiple modules might need a different theory to explain". While I understand this is a big picture question that maybe isn't entirely fair to ask of the authors, I think: 1) the authors do a nice job of positioning their review as a kind of discussion on what normative modeling can provide to neuroscience, so having this discussion on when the failure of a model to capture ALL aspects of the biological features motivates further constraints as opposed to a new approach, would be useful; 2) this question connects with the title of the paper, i.e. "what is the question?"

    4. Reviewer #3 (Public review):

      Summary:

      The authors present an extensive review of the literature on normative grid cell theory, asking what kind of cost function might be minimized by the entorhinal grid cell code. The authors show which of the main features of grid cells emerge from combinations of terms in a cost function that optimizes for spatial fidelity, biological plausibility, and path integration. They conclude by outlining potential future directions for the field.

      Strengths:

      The structure of the review makes it particularly useful for researchers who are familiar with grid cells but not necessarily with normative models. Equations are kept to a minimum and are usually explained conceptually.

      Weaknesses:

      I identified one main weakness, related to the fact that the introduction to experimental results around grid cells and what they allow us to conclude is less nuanced than the rest of the review. However, since this is not the main focus of the manuscript, I consider this a secondary limitation.

      The review organizes the current literature on the subject within a coherent conceptual framework, helping to define possible paths forward for the field.

    5. Author response:

      We thank the reviewers for their time and attention which will significantly improve the paper. Further, we are grateful for their appreciation of our goals and work. In sum, the reviewers point to our overstated discussion of experimental evidence which we will tone down, some slightly confusing points of argumentation which we will clarify, and some discussion points on the role of normative theories that we will add text to address. We believe this will improve the paper significantly and hope you agree!

      Major Concern: Experimental Support for Path-Integration is not as strong as suggested

      The major point raised by all reviewers (reviewer 1 comment 1, reviewer 2 comment 1, reviewer 3’s only weakness) was that our presentation of the experimental perturbation evidence for path-integration is stronger than the reality. On reflection, we agree with this evaluation. We thank the reviewers for raising it; we will moderate our writing and include the sensible caveats raised. In sum, we still think that the convergence of evidence points to path-integration: first, disruptions to grid cells lead to path-integration problems, though these perturbations admittedly aren’t perfectly precise; second, normative theories of path-integration lead to grid cells and predict grid cell behaviour; third, mechanistic models of path-integration match grid cell behaviour and predict connectivity subsequently measured in entorhinal cortex. However, the evidence is not as all-encompassing as we suggested.

      That said, we’d like to further comment on one point. It is argued (reviewer 1, comment 1) that there are other theories of grid cell function, and that we discuss these theories. We discuss efficient-coding only models of grid cells and emphasise strongly why we reject them. We also briefly discuss oscillatory-interference models of path-integration and our reasons for not pursuing them further. As such, the reviewer is correct that our reading of literature strongly points us towards path-integration rather than other theories. We will slightly change the framing of the paper to make it clear that we are making a case. However, we are not aware of other theories the reviewer might be referring to. If the reviewer can point us to the other suggested theories that we do not address we would be happy to evaluate and include them.

      We now turn to the remaining comments, and how we plan to address them.

      Reviewer 1, Comment 2 – There could be multiple roles for grid cells

      The reviewer is indeed right that grid cells might perform multiple functions. This could just mean that the same computational motif (e.g. path-integration) is reused across different computations though that introduces no changes to the required normative theory. A stronger claim would be that grid cells perform both path-integration and some other function. This, according to a normative perspective, would most likely change how grid cells were optimally structured. We use the fact that large parts of the grid cell code can be captured with only path-integration as an argument against additional roles for grid cells. That said, there exist properties of grid cells not well-captured by path-integration which could well be smoking guns for additional roles of grid cells. The review already discusses both discrepancies between grid cells in three and two dimensions, and inhomogeneities in the grid in complex environments, and we will add two more (heading direction and peak-to-peak/angular variability, discussed below) that we are grateful to the reviewers for raising, and we discuss each of these in detail below.

      That said, whether these are necessarily arguments against purely path-integration or a reflection of interesting mappings of the core path-integration mechanism to the measurements we make remains to be seen. We would argue that both 3D grid cells (as explained below: there appear to be 2D slices in which grid cells behave as you’d expect) and spatial inhomogeneities (as explained in the paper: mappings of torus to world can introduce warping) can be explained without reference to additional computational roles of grid cells, which remain to us the most parsimonious explanation. We discuss next the slight update to path-integration only that the heading direction story suggest. But in sum, our view is that these discrepancies are likely not fatal for our path-integration-centric view of grid cells, but may well suggest some very interesting clarifications.

      Reviewer 1, Comment 4 – The system has two heading signals: true & internal, why?

      The reviewer is right to point to the puzzle over true vs. purely internal heading direction and which drives grid cells. We believe recent work from Abraham Vollan has effectively solved this puzzle: there appear to be two parallel circuits, one theta-modulated and following internal heading direction, another theta-unmodulated and aligning more with true heading direction. We will make sure to include discussion of this exciting work in our revised submission. This serves as a good example of an update we concede to the most austere version of the path-integration only view. Rather, it seems there are two parallel path-integrators working with different heading signals. The reasons for this remain unclear, but seem to be related to attention and planning (Vollan et al. 2026).

      Reviewer 2, Comment 3: Real Grid Cells have peak-to-peak variability & Angular variability

      The reviewer is right to point to the discrepancy in peak-to-peak firing rate and angles within a module that we did not adequately address. First, it is Sorscher’s RNN models, not nonnegative PCA that can generate a distribution of grid angles (Redman et al. 2025), which suggests that path-integration and such variability are compatible. We emphasise this point because the non-path-integration results from nonnegative PCA produce grid cells oriented at 30 degree offsets, something not measured even when you’re careful as in Redman et al. 2025. Thus, this becomes an interesting target for future work: perhaps using theories of path-integration up to an error threshold (rather than perfect) such angular diversity would be recovered. We will include this in our discussion. Further, we will include discussion of peak-to-peak variability that, as yet, has no obvious role.

      Reviewer 2, Comment 1: grid cells are inhomogeneous in 3D or complex environments, doesn’t that break the theory?

      Disrupted grid coding in extended or 3D environments indeed deserve more discussion, which we will add. In particular, we will add recent evidence that grid cells in 3D can be understood via the correct sequence of 2D projections(Qi & Yartsev, 2026). These two phenomena seem, to us, consistent with a path-integration only view of grid cells, as discussed above, and we hope to make this position clearer.

      Reviewer 2, Comment 5: Couldn’t there be other reasons for multiple modules?

      We have suggested a consistent normative framework in which multiple modules are explained through their role in non-linear coding. We think this elegant, and the most parsimonious current theory. We could, of course, be wrong. The discrepancies pointed to above might be good clues to follow to work out what else these modules might be doing, but currently these alternative explanations seem not to exist. We will text to clarify this.

      Reviewer 1, Comment 3: The review confuses computational and parameter parts of normative theory

      We disagree with the reviewer’s dichotomisation of normative theory. We view a normative theory as the complete procedure that produces the predictions. Almost all such theories have parameters and hence fitting a theory to data comprises both elements (a) [computational role] and (b) [specific parameters] identified by the reviewer. Occasionally theories have no parameters in the traditional sense, e.g. Rebecca et al.; instead they have heavy assumptions that play an equivalent role. It is true that, as the reviewer says, Sorscher et al.’s work was criticised for producing grid cells only for specific parameter values. We never found this as damning as Schaeffer et al. argued: simply it says that that theory is only correct within the given parameter range. Rather, arbitrating between models, parameters, or assumptions seems the same basic process: see what they predict and keep working with models while they remain useful ways to understand measured phenomena. If a model with very specific parameter values remains useful, that seems okay. In fact, we argued extensively why we think the nonnegative PCA model is not a useful model, but this was for completely different reasons. To us this story just reinforces the importance of hygiene in normative research: perform parameter sweeps and clarify how they constrain the claims you are making, carefully arbitrate what models can capture. Indeed, that is the whole goal of this review. We might be misunderstanding and, if so, we welcome correction.

      Reviewer 2, Comment 4: Normative Models of Cells Beyond Grid Cells

      The reviewer is right that extending these models to other cell types is an interesting area for further work, and that other cell types do seem to be involved in aspects of navigational computations both in RNNs and the brain. We will include a discussion to this effect in the revised manuscript. That said, we think the modularity of grid cells and their tight-linking to path-integration calculations should also be appreciated as a win!

      Reviewer 2, Comment 2: Multi-modularity is not cleanly explained

      We thank the reviewer for the comments, we agree. We will clarify the story regarding multiple modules, and will explain the equation further.

      Reviewer 1, Comment 5: the early introduction of phase-shifted Grid Cells seem the perfect place to normatively argue for Path-integration!

      We agree with the reviewer that this point can be made both normatively (‘oh look! If I try to do this optimally, I get translations!’) or, as we did early in the paper, mechanistically (‘oh look! With these cells I can do this!’). Indeed, a large part of the point of our paper is that path-integration is what is required to normatively derive phase-shifted grid modules, something discussed by Rebecca et al., our earlier work, and RNN studies, and appreciated for two decades. The earlier part of the paper does not discuss these papers as that section is aimed at giving intuition for the solution (mechanism). Later sections then heavily discuss the normative angle. We hope that division of labour makes sense.

      Finally, we will refine our summary of Rebecca et al. The reviewer is right that neurons don’t have to be discrete, we apologise for that error, but our understanding is that the only meaningful role of a neuron in Rebecca et al.’s work is the region in which is active, effectively making every neuron a binary unit, which seems dubious. We will clarify that by “predict velocity from each current and next encoding” we mean that the normative constraint they enforce is axiom 1: sequential activity of sets of neurons i then j can be uniquely interpreted as a trajectory, i.e. a step or velocity. Their work is elegant, and we will try to do more justice to it in the revision.

      To conclude, we thank the reviewers for their extensive comments, and look forward to releasing a version that addresses their concerns.

    1. eLife Assessment

      This study presents a valuable and well-documented computational pipeline for the scalable analysis and spike sorting of large extracellular electrophysiology datasets, with particular relevance for high-density recordings such as Neuropixels. The authors demonstrate the pipeline's utility for benchmarking spike sorter performance and evaluating the effects of data compression, supported by thorough testing, clear figures, and openly available code. The workflow is reproducible, portable, and practical, providing concrete guidance on computational cost and runtime. Overall, the evidence supporting the pipeline's performance and output quality is compelling, and this work will be of broad interest to the systems neuroscience community.

    2. Reviewer #1 (Public review):

      Summary:

      Extracellular electrophysiology datasets are growing in both number and size, and recordings with thousands of sites per animal are now commonplace. Analyzing these datasets to extract the activity of single neurons (spike sorting) is challenging: signal to noise is low, the analysis is computationally expensive, and small changes in analysis parameters and code can alter the output. The authors address the problem of volume by packaging the well-characterized SpikeInterface pipeline in a framework that can distribute individual sorting jobs across many workers in a compute cluster or cloud environment. Reproducibility is ensured by running containerized versions of the processing components.

      The authors apply the pipeline in two important examples. The first is a thorough study comparing the performance of two widely used spike-sorting algorithms (Kilosort 2.5 and Kilosort 4). They use hybrid datasets created by injecting measured spike waveforms (templates) into existing recordings, adjusting those waveforms according to the measured drift in the recording. These hybrid ground truth datasets preserve the complex noise and background of the original recording. Similar to the original Kilosort 4 paper, which uses a different method for creating ground truth datasets that include drift, the authors find Kilosort 4 significantly outperforms Kilosort 2.5. The second example measures the impact of compression of raw data on spike sorting with Kilosort 4, showing that accuracy, precision, and recall of the ground truth units is not significantly impacted even by lossy compression. As important as the individual results, these studies provide good models for measuring the impact of particular processing steps on the output of spike sorting.

      Strengths:

      The pipeline uses the Nextflow framework, which makes it adaptable to different job schedulers and environments. The high-level documentation is useful, and the GitHub code is well organized. The two example studies are thorough and well-designed and address important questions in the analysis of extracellular electrophysiology data.

      Weaknesses:

      There are no major weaknesses in the revised manuscript. While no data analysis pipeline can cover the needs of all experiments, the authors have added and significant flexibility in the pipeline. Even experimenters who might opt for a simpler pipeline will benefit from this work as a model.

    3. Reviewer #2 (Public review):

      Summary:

      This work presents a reproducible, scalable workflow for spike sorting that leverages parallelization to handle large neural recording datasets. The authors introduce both a processing pipeline and a benchmarking framework that can run across different computing environments (workstations, HPC clusters, cloud). Key findings include demonstrating that Kilosort4 outperforms Kilosort2.5 and that 7× lossy compression has minimal impact on spike sorting performance while substantially reducing storage costs.

      Strengths:

      (1)Extremely high-quality figures with clear captions that effectively communicate complex workflow information.

      (2) Very detailed, well-written methods section providing thorough documentation.

      (3) Strong focus on reproducibility, scalability, modularity, and portability using established technologies (Nextflow, SpikeInterface, Code Ocean)

      (4) Pipeline publicly available on GitHub with documentation.

      (5) Clear cost analysis showing ~$5/hour for AWS processing with transparent breakdown.

      (6) Good overview of previous spike sorting benchmarking attempts in the introduction

      (7) Practical value for the community by lowering barriers to processing large datasets.

      Weaknesses

      No significant weaknesses. The authors have responded to all my review critiques and suggestions.

    4. Reviewer #3 (Public review):

      Summary:

      The authors provide a highly valuable and thoroughly documented pipeline to accelerate the processing and spike sorting of high-density electrophysiology data, particularly from Neuropixels probes. The scale of data collection is increasing across the field, and processing times and data storage are a growing concern. This pipeline provides parallelization and benchmarking of performance after data compression that helps address these concerns. The authors also use their pipeline to benchmark different spike sorting algorithms, providing useful evidence that Kilosort4 performs the best of out the tested options. This work, and the ability to implement this pipeline with minimal effort to standardize and speed up data processing across the field, will be of great interest to many researchers in systems neuroscience.

      Strengths:

      The paper is very well written and clear. The accompanying GitHub and ReadTheDocs are well organized and thorough. Benchmarks are exceptionally well applied to support the authors' claims, and it is clear that the pipeline has been very thoroughly tested and optimized by users at the Allen Institute for Neural Dynamics. The pipeline incorporates existing software and platforms that have also been thoroughly tested (such as SpikeInterface), so the authors are not reinventing the wheel, but rather putting together the best of many worlds. In the latest revision, the authors add a nice analysis showing that compression mostly affects the lowest SNR units. This is a great contribution to the field and it is clear the authors have put a lot of thought into making the pipeline as accessible as possible.

      Weaknesses:

      None noted. The authors have addressed all previous questions and requests for clarification.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Weaknesses:

      The pipeline is very complete, but also complex. Workflows (optimal artifact removal, best curation for data from a particular brain area or species) will vary according to experiment. Therefore, a discussion of the adaptability of the pipeline in the “Limitations” section would be helpful for readers.

      We added a dedicated paragraph in the Discussion section under “Limitations” focusing explicitly on the adaptability and flexibility of the pipeline. Furthermore, we took this feedback as an opportunity to make the pipeline itself significantly more modular and customizable with the most recent release (v1.2.0: https://aind-ephys-pipeline.readthedocs.io/en/latest/releases/1.2.0.html).

      Reviewer #1 (Recommendations for the authors):

      (1) In the description of the Phase-shift correction (Line 166-167): The current text reads “As a result, different groups of channels are sampled asynchronously.” A better description would be: “Sample times for different groups of channels are offset in time by a known amount.”

      We replaced the phrase in the manuscript text with the suggested formulation.

      (2) Figure 5 and description of the benchmarking overview (Line 326-336): How were spike trains (times) selected for the injected ground truth units? What was the range of firing rates?

      All injected spike trains were generated as independent Poisson processes featuring a mean firing rate of 15 Hz. We have now incorporated this explicitly into the main text to clarify the ground-truth injection process.

      (3) Figure 6, panel b: Are the gray points in the raster the original spikes in the test recording? From the pattern, it looks like there are 8 recovered ground truth units. Were the other 2 undetected by either sorter?

      That is correct; the two remaining units were undetected by both sorters. To clear up any confusion, we updated the caption for Figure 6 to state: “Note that spikes undetected by any of the sorter are not shown in the plot.”

      (4) Figure 7, panel c: Are all units returned from KS included in these distributions? (i.e., regardless of the KS refractory metric calculated by the sorter) - it would be useful to add that detail to the caption. It would also be helpful for panel C to include a total unit count from the two sorters... Also, since there are multiple ways to calculate the refractory period contamination, it would be good to state the calculation used here.

      Because we rely directly on the hybrid ground-truth for accurate validation, we included all raw units returned by Kilosort for this specific analysis. We have explicitly added a note detailing this to the caption. Panel C does report the total raw unit count returned by the two sorters (N = 3046 for KS2.5; N = 3652 for KS4).

      Additionally, to clarify the evaluation procedure, we appended the following statement to the main text: “For all results, we perform spike train comparisons and compute performance metrics as defined in (Buccino et al. 2020), using all units returned by the spike sorter (without any sorterspecific curation).”

      (5) Comments about the pipeline:

      The paper clearly demonstrates the immense utility of the pipeline in the authors’ work. I did some testing to try to understand its adaptability to workflows at my institution.

      I tested the pipeline on our local cluster running LSF. I’ve worked on a similar pipeline using Nextflow to automate ephys analysis with the same sorters. Questions that came up for me that would be usefully addressed in the ’Limitations’ section:

      (i) Is the pipeline meant to be run only in total? In particular, is it possible to start with preprocesseddata? (aind-ephys-preprocessing/code/params.json does not appear to include any means to turn off filtering, for example). Is the pipeline meant to be run only in total? In particular, is it possible to start with preprocessed data? (aind-ephys-preprocessing/code/params.json does not appear to include any means to turn off filtering, for example).

      To accommodate users who wish to run only parts of the workflow or use external preprocessing setups, we have refactored the codebase to support a custom preprocessing pipeline option. This makes it possible to turn off standard filtering or inject custom workflows.

      (ii) For debugging purposes, is there a means to go from preprocessing or sorting to result collection,so that interim results can be interpreted even when some steps of the pipeline aren’t working?

      The pipeline is designed to be a spike sorting pipeline, so the spike sorting step cannot be skipped. However, we have rewritten the post-sorting architecture to make it highly lightweight and fault-tolerant. The postprocessing step now only requires the random spikes and templates computation and downstream steps have been update to accomodate this lightweight option. As an example, if no quality metrics are computed, the curation step will be skipped. The visualization and QC steps also required updates to be tolerant to missing extensions. This required coordinate updates across several components:

      Postprocessing: PR #12

      Curation: PR #13

      Visualization: PR #21

      Quality Control: PR #20

      (iii) If these options to skip processes and output data ’partway’ are available, it would be great toadd that to the documentation.

      We have fully updated our online documentation for v1.2.0 (release notes: https://aind-ephys-pipeline.readthedocs.io/en/latest/releases/1.2.0.html), introducing a brandnew “Customization” guide page that comprehensively explains how to construct and provide custom preprocessing and postprocessing strategies, as well as how to integrate a new spike sorter in the pipeline: https://aind-ephys-pipeline.readthedocs.io/en/latest/customization.html

      Reviewer #2 (Public review):

      Summary:

      This work presents a reproducible, scalable workflow for spike sorting that leverages parallelization to handle large neural recording datasets. The authors introduce both a processing pipeline and a benchmarking framework that can run across different computing environments (workstations, HPC clusters, cloud). Key findings include demonstrating that Kilosort4 outperforms Kilosort2.5 and that 7× lossy compression has minimal impact on spike sorting performance while substantially reducing storage costs.

      Strengths:

      (1) Extremely high-quality figures with clear captions that effectively communicate complex workflow information.

      (2) Very detailed, well-written methods section providing thorough documentation.

      (3) Strong focus on reproducibility, scalability, modularity, and portability using established technologies (Nextflow, SpikeInterface, Code Ocean).

      (4) Pipeline publicly available on GitHub with documentation.

      (5) Clear cost analysis showing ~$5/hour for AWS processing with transparent breakdown.

      (6) Good overview of previous spike sorting benchmarking attempts in the introduction.

      (7) Practical value for the community by lowering barriers to processing large datasets.

      Weaknesses:

      No significant weaknesses were identified, although it is noted that the limitations section of the discussion could be expanded.

      We thank the reviewer for their constructive feedback on our manuscript.

      Reviewer #2 (Recommendations for the authors):

      The authors could discuss why 2.25 bps is the “lowest supported” level and whether more aggressive compression could be achieved with custom approaches, potentially exploring where performance breakdown occurs.

      The 2.25 bits-per-sample (bps) limit is an inherent constraint of the WavPack lossy compression library itself. While more aggressive, domain-specific, or custom compression schemes could be explored, we focused on WavPack due to its native support in modern neurophysiology ecosystems and its excellent performance in our prior simulated benchmarks (Buccino et al. 2023). We agree that using this hybrid benchmarking framework to explore alternative compression configurations is a highly valuable avenue for future work. We have added the following text to the Discussion: “The benchmarking pipeline will continue to develop as an open evaluation framework, enabling transparent and reproducible comparisons of spike sorting and preprocessing methods across the community. As one example, the work on lossy compression could be extended with additional codecs and parameter settings, exploiting our ability to read out spike sorting degradation directly from the hybrid ground truth spike times.”

      (2) The limitations section would benefit from expansion to include: (i) discussion of how simulated data limitations may affect generalization of benchmarking results to real neural data, and (ii) clarification of the effort required to add new spike sorters, including configuration complexities for coordinating Nextflow processes beyond simple SpikeInterface integration.

      We have expanded the Discussion section to address both items:

      (i) We added a paragraph detailing the specific limitations of hybrid ground-truth datasets (e.g., how idealized template injection might miss extreme multi-unit overlapping dynamics or nonstationary noise properties found in real tissue).

      (ii) We added a structural overview section clarifying the workflow complexity, detailing exactly what steps are required to map a new spike sorter into a Nextflow execution processes beyond its baseline addition to Spike Interface.

      (3) The authors should clarify the terminology of “hypothetical experiment” in the introduction to improve reader comprehension.

      We have removed the word hypothetical from the introduction to ground the explanation more directly.

      (4) The cost analysis could be improved by making it clearer whether “runtime” refers to wall-clock vs. total parallel compute time.

      We mean wall-clock time. While total parallel compute time aggregated across cloud workers remains roughly identical to the overall sequential execution on a lone cloud instance, cluster parallelization slashes the wall-clock time drastically. We have updated the text to explicitly state that reported runtimes represent wall-clock time.

      (5) The authors could address the Nextflow Java dependency limitation by discussing containerized execution options (Docker/Singularity) as a solution, while noting relevant HPC system restrictions.

      We have updated the text to mention the official pre-built Nextflow container images as an elegant workaround for environments where local Java installations are blocked or restricted: “However, one option to bypass installation issues is to run the main pipeline script in container images packaged with Nextflow (https://hub.docker.com/r/nextflow/nextflow).”

      (6) Figure 8 analysis would be strengthened by explicitly noting that compression effects are more substantial for lower-accuracy units, suggesting better preservation of higher SNR units.

      We appreciate this insight. To evaluate this systematically, we generated a new supplementary figure (Figure S3) which shows sorting performance during lossy compression as a function of the Signal-to-Noise Ratio (SNR) of ground truth units. The plot demonstrates that for Neuropixels 2.0 recordings, the slight drop in sorting accuracy is indeed heavily concentrated among low-SNR units. We have integrated this observation into the Results section.

      Reviewer #3 (Public review):

      (1) Could the authors please expand on the statement on line 274, that processing their test dataset serially “on a single GPU-capable cloud workstation... would take approximately 75 hours and cost over 90 USD.” How were these values calculated? I was a bit surprised that this is a ¿4-fold slowdown from their pipeline, but only increases the cost by 1.35x... More context on why this is, and maybe some context on what a g4dn.4xlarge is compared to the other instances, might help.

      We have expanded the cost analysis section in the manuscript methods to explain these figures explicitly. The serial run relies on a single continuous, higher-tier GPU workstation instance (g4dn.4xlarge) running uninterrupted for 75 hours.

      Our distributed pipeline, by contrast, dynamically provisions CPU-only instances to process chunked preprocessing steps concurrently, then spins up short-lived GPU spot instances only when Kilosort executes. While this parallel execution compresses the overall wall-clock time by over 4-fold, the cost is only moderately reduced because the CPU-only instances with many parallel processing cores are only slightly less expensive than GPU instances.

      (2) One of the most commonly used preprocessing pipelines for Neuropixels data is the CatGT/ecephys pipeline from the developers of SpikeGLX at Janelia. It may be worth commenting very briefly... on how the preprocessing steps available in this pipeline compare to the steps available in CatGT. For example, is “destriping” similar to the “-gfix” option in catGT to remove high-amplitude artifacts?

      We have added a section drawing direct comparisons to CatGT preprocessing workflows. We explicitly clarify that our phase-shift correction performs the exact same function as CatGT’s Tshift. We also point out that while our current version lacks a direct equivalent to CatGT’s saturation removal feature (-gfix), this capability is scheduled for incorporation in our upcoming pipeline release.

      (3) Why are there duplicate units (line 194), and how often is this an issue? I understand that this is likely more of a spike sorter issue than an issue with this pipeline, but 1-2 sentences elaborating why might be helpful for readers.

      Duplicate units are primarily an artifact of template-matching sorting routines (such as Kilosort), which can occasionally split a single biological neuron into multiple overlapping spatial templates or over-extract templates in highly active channel regions. We have added two clarifying sentences explaining this phenomenon in the text: “Next, duplicated units, that can arise when using template-matching methods if different templates are consistently fit to the same spikes, are removed based on the fraction of overlapping spikes.”

      Customizability of cluster curation parameters It seems from the parameter files on GitHub that the cluster curation parameters are customizable - correct? If so, it may be worth explicitly saying so in the curation section of the text... A presence ratio of >0.8 could be particularly problematic for some recordings (e.g. state transitions, behavior specific cells).

      (4) Yes, they are completely customizable. We agree that a rigid presence ratio cutoff of 0.8 would erroneously discard highly valid units that are modulated by specific behavioral states, or are active only during sleep vs. wake cycles. We have explicitly added text in the Curation section clarifying that all quality metric thresholds can be modified by the user: “Units are tagged as passing a default_qc when they satisfy the following criteria based on quality metrics thresholds. Thresholds can be user defined, and these are the default”.

      (5) The axis labels in Figures 3d-e are too small to see, and Figure 3d would benefit from a brief description of what is shown.

      We have updated the figures with enlarged, high-visibility axis labels and expanded the caption of Figure 3d to clearly describe the visualization.

      Figure 4 labels (“neural” vs “passing QC”) (6) What is the difference between “neural” and “passing QC” in Figure 4?

      We have updated the figure caption for Figure 4 to include an explicit cross-reference to the Curation methodology section, which defines the strict quantitative boundary between raw neural classification and formal automated QC passage.

      (7) I understand the current paper is focused on spike data... but I am curious about the NP2.0 probes that save data in wideband. Does the lossy compression negatively affect the LFP data? Is software filtering applied for the spike band before or after compression?

      Compression is applied to the raw streams prior to any secondary downstream software processing. For Neuropixels 1.0, compression is executed strictly on the action potential (AP) stream. For Neuropixels 2.0, compression operates directly on the unified wide-band data stream.

      Software filtering to separate bands is conducted post-decompression, as captured in our baseline workflow definitions (e.g., WavPack compression → decompression → preprocessing → Kilosort4). To clarify this, we added the following text: “In all cases, compression was applied before any preprocessing took place. For Neuropixels 1.0, we compressed the AP stream only. For Neuropixels 2.0, we compressed the full wide-band data.”

      Because LFP signals possess inherently smooth continuous dynamics across both space and time, they are much more amenable to lossless or near-lossless compression. Thus, the minor losses introduced by lossy compression are overwhelmingly localized to high-frequency spike band features, leaving LFP components virtually unaffected.

    1. eLife Assessment

      The authors describe a valuable finding that the Streptococcus pyogenes secreted protease SpeB is expressed in response to protease activity that degrades the Vfr repressor. Proteases can be released from host neutrophils (possibly by NETosis), as well as a positive feedback mechanism by SpeB itself. The authors utilize a dual fluorescent reporter system to simultaneously read speB and capsule gene expression, providing solid evidence that demonstrates that proteases can regulate Vfr; however, the data indicating that this is physiologically relevant and that extracellular traps themselves have a functional role are incomplete. This work will be of interest to microbiologists studying the regulation of virulence factors at the host-pathogen interface.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript examines how Streptococcus pyogenes regulates expression of the virulence factor SpeB in response to both bacterial and host-derived cues. The authors propose that Vfr acts as a repressor of speB expression and that degradation of Vfr by SpeB or by neutrophil-derived proteases relieves this repression. This creates a model in which S. pyogenes can sense proteolytic activity during infection and use that information to tune virulence factor expression.

      Strengths:

      The main strength of the study is the bacterial regulatory mechanism. The dual reporter system provides a useful way to follow speB and hasABC expression, and the genetic analysis of known regulators helps validate the system. The media-swap experiments, recombinant Vfr experiments, and SpeB-mediated degradation of Vfr support the conclusion that Vfr represses speB and that proteolysis can relieve this repression. The finding that SpeB can degrade Vfr is particularly interesting because it suggests an autoregulatory mechanism that could reinforce SpeB expression once it has been initiated.

      Weaknesses:

      The host side of the model is less completely supported. The authors show that neutrophil lysates and protease-containing fractions can induce the speB reporter and degrade Vfr, which supports the idea that neutrophil-derived proteases can affect this circuit. However, the in vivo interpretation relies heavily on PAD4-deficient mice to implicate neutrophil extracellular traps. PAD4 deficiency is a useful perturbation, but it does not by itself distinguish loss of extracellular trap formation from changes in neutrophil recruitment, survival, degranulation, phagocytosis, oxidative killing, or other neutrophil death pathways. As a result, the current data support a role for neutrophil-associated proteolytic activity more strongly than they support a specific role for extracellular traps. This distinction is important for interpreting the central model. The bacterial circuit is well developed, but the host-derived cue remains somewhat underdefined. If the relevant signal is extracellular protease activity more broadly, then the model is still interesting, but the conclusion should be framed around neutrophil-derived proteolytic stress rather than extracellular traps specifically. If extracellular traps are the key in vivo source of protease exposure, then additional evidence would be needed to separate that mechanism from other neutrophil effector functions that remain intact in PAD4-deficient cells.

      Overall:

      This is a valuable study with solid evidence for a bacterial protease-sensing regulatory mechanism controlling SpeB expression. The work should be useful to investigators interested in bacterial virulence regulation, host-pathogen interactions, and how pathogens integrate immune-derived cues during infection. The impact of the study would be stronger if the host-derived signal were defined more precisely, but the bacterial Vfr-SpeB circuit provides a compelling framework for thinking about how S. pyogenes links proteolytic activity to virulence gene expression.

    3. Reviewer #2 (Public review):

      Summary:

      The study examines how Streptococcus pyogenes integrates bacterial and host-derived signals to regulate SpeB, proposing that Vfr acts as a protease-sensitive repressor whose degradation relieves repression of speB. The authors further suggest that neutrophil-derived serine proteases, including those associated with inflammatory conditions, may promote this transition, and thereby counterbalance LL-37/CovRS-associated suppression of speB. The conceptual framework is interesting and potentially important for understanding how host inflammation feeds into bacterial virulence regulation.

      Strengths:

      The work addresses a biologically significant question and does so using a broad and generally well-integrated experimental approach, including bacterial genetics, reporter assays, recombinant protein analyses, neutrophil-derived material, human blood infection, and mouse infection models. A particular strength is the effort to connect host inflammatory processes to bacterial regulatory behavior, which gives the study conceptual reach beyond a narrow mechanistic observation. The data support the view that Vfr is relevant to speB control and that neutrophil-associated protease activity may influence this pathway.

      Weaknesses:

      The main limitations are mechanistic. The physiological form, localization, and abundance of Vfr are not sufficiently defined to support the proposed model at full strength, and the evidence that Vfr functions as a SpeB-labile repressor under biologically relevant conditions remains incomplete. The relationship between Vfr and the broader RopB/SIP regulatory framework is also not yet firmly established. In addition, the reporter system is not yet benchmarked closely enough against endogenous SpeB protein output, and its growth-phase dependence is insufficiently resolved, which makes it difficult in some settings to distinguish promoter activity from mature protease production. The neutrophil protease component is likewise not defined beyond a general serine protease signal, and the potentially important LL-37/CovRS/Vfr connection is underdeveloped in the main text. Overall, the conceptual advance is promising, but several of the central mechanistic claims would benefit from more direct experimental support and more cautious framing.

    4. Reviewer #3 (Public review):

      Summary:

      SpeB is a cysteine protease secreted during infection by Streptococcus pyogenes (Spy). SpeB has been extensively investigated for its role in pathogenesis, which involves proteolytic processing of both Spy virulence factors and host proteins. Regulation of speB expression is complex and includes growth phase regulation, a quorum-sensing system, the transcription factor RopB, and the global regulatory system CovRS (CsrRS). Guerra et al now attempt to refine the current model of regulation of SpeB expression, focusing on the Spy protein Vfr, which has been suggested previously to act as a negative regulator of SpeB expression. In the current study, neutrophil lysates (representing proteases released during NETosis) are shown to degrade Vfr and to relieve repression of SpeB. At high cell density, SpeB itself also degrades Vfr, which may allow autoregulation of SpeB expression. These observations are unsurprising as the broad protease activities of both neutrophil proteases and SpeB are well known. Nonetheless, the data presented fill in additional details in our understanding of the complex regulation of an important Spy virulence factor.

      Strengths:

      (1) Construction of a GFP reporter strain provided a facile methodology for tracking speB promoter activity in a variety of experimental setups.

      (2) A Vfr deletion mutant was a useful tool to investigate the role of Vfr in SpeB regulation, and mutants in speB and ropB were important controls.

      (3) Experiments using neutrophil lysates in vitro, as well as in vivo studies of mice depleted of neutrophils with anti-Ly6G or in PAD4-/- mice (that cannot form NETs) support the hypothesis that neutrophil proteases derepress speB expression by degrading Vfr.

      Weaknesses:

      (1) The introduction and all the experiments in Figure 1 focus on CovRS, which turns out to be largely tangential to the overall story developed by the rest of the study. On the other hand, the complex and well-studied regulation of speB expression by RopB and the SIP quorum-sensing system is only minimally described. A better framing would be a more detailed introduction to the current model of speB/RopB/SIP/quorum sensing/growth phase regulation. CovRS could be introduced later as its relevance is really just to show that neutrophil lysates or NETs do more than simply providing LL-37, which signals through CsrS, as another regulator of speB expression.

      (2) Vfr, as the central focus of the paper, also deserves a more thorough introduction to provide context for the study. For example, reference 19 (Shelburne et al, 2011) showed reduced transcription of speB in a vfr mutant, an effect that could be complemented by expressing vfr or a 39-aa N-terminal fragment in trans. That study presented evidence that the N-terminal peptide binds to RopB, which may prevent RopB from upregulating SpeB expression. Do the authors concur with that model? As it stands, the discussion and model in Figure 1A imply a direct regulatory effect of Vfr on speB expression rather than an indirect one through regulation of RopB. If direct regulation of speB by Vfr is a consideration, it should be investigated more thoroughly, e.g., by promoter-binding assays, CHIP-seq, etc.

      (3) Use of single-cell flow cytometry generally confirmed results observed in batch culture. The authors also comment repeatedly on the heterogeneity of individual cell fluorescence representing both speB and has operon expression. However, the reason(s) for heterogeneity in gene expression are not explored, e.g., differences in individual cell growth rate in batch culture, variable loss of reporter plasmid during infection experiments, etc).

      (4) Lines 116-118 and Figure 3C: Incubation of recombinant Vfr with Spy Dvfr reduced SpeB expression, but the degree of suppression is modest compared to that seen in wild-type Spy. How does the concentration of rVfr added compare to that present in the culture fluid of wild-type Spy? (Also, the concentration of rVfr used is unclear: the figure says 3 µg/ml and the legend says 0.3 mg/ml, i.e., 300 µg/ml).

      (5) Lines 125-126: "...the Vfr structure contains several potential protease SpeB cleavage sites..." The role of Vfr in degrading SpeB could be clarified by identifying the predicted cleavage products, e.g., by mass spec, after co-incubation of the two recombinant proteins.

      (6) Lines 122-124: "Notably, speB expression in Spy Dvfr is unaffected by LL-37 or MgCl2, further validating its [Vfr's?] dominance over CovRS regulation." This statement is an oversimplification and is potentially misleading: LL-37 is degraded by SpeB (Nyberg et al, JBC 2004), which likely explains why the addition of LL-37 fails to signal through CovRS to repress SpeB in Spy Dvfr since SpeB is produced continuously in that strain. By contrast, SpeB is only produced during the stationary phase in the wild type, so LL-37 remains active throughout the exponential phase and represses SpeB expression. The response to the CovRS ligand MgCl2 is similar (or greater) in Spy Dvfr compared to wild type (Figure S2C).

      (7) Lines 153-154 and Figure 6E: Growing wild type Spy in the presence of neutrophil lysates with or without a protease inhibitor stimulated or repressed speB expression in a manner consistent with degradation (or not) of Vfr. It would be confirmatory and informative to do the same experiment with the Spy Dvfr strain.

      (8) Clarity of writing could be improved, particularly by eliminating pronouns of indefinite reference (it, its, this) in contexts in which the subject is ambiguous (examples at lines 62, 89, 111, 114, 115, 123, 183, 190, 193, 204, 205, 210, 217, 221, 222, 224).

    1. eLife Assessment

      The study presents valuable findings of a new E. coli cell-free protein synthesis (eCFPS) system that has been simplified by reducing the number of core components from 35 to 7; furthermore, the findings communicate a simplified 'fast lysate' preparation that eliminates the need for traditional runoff and dialysis steps. It is interesting that the system's robustness is exhibited by its applicability to nanoluc, a protein that expresses readily in many systems, to more challenging proteins like the functional self-assembling vimentin and the active restriction endonuclease Bsal. Despite the study representing an advancement towards simplifying protein expression workflows, the evidence is solid and supports the main claims however minor weakness exists i.e. the efficiency claims about the new system needs to be supported by accurate comparisons with typical cell free expression systems, in addition, investigations into the mechanistic basis of the observations would provide more evidence. Despite this shortcoming, the paper remains of interest to scientists in cell and molecular biology, microbiology, biotechnology and protein synthesis.

    2. Reviewer #1 (Public review):

      Summary:

      The authors presented a simplified E. coli cell-free protein synthesis (eCFPS) system reduces core reaction components from 35 to 7, improving protein expression levels. They also presented a "fast lysate" protocol that simplifies extract preparation, enhancing accessibility and robustness for diverse applications.

      Strengths:

      The authors present a valuable new protocol for eCFPS, which simplifies its application.

      Weaknesses:

      The authors provide data for optimization but offer insufficient explanation of the fundamental mechanisms underlying the phenomenon based on data.

      Comments on revised version.

      The authors have satisfactorily addressed the concerns raised by the reviewers. However, the mechanistic basis of the observed performance gain remains insufficiently substantiated. The attribution of this improvement to enhanced transcription is currently speculative. This point could be directly tested by quantifying mRNA levels, for example, using real-time PCR, in both the initial and optimized systems. Such analysis would significantly strengthen the mechanistic interpretation of the results.

    3. Reviewer #2 (Public review):

      Summary:

      The authors have made a convincing argument that the current system of in vitro translation using E. coli extracts can be significantly optimized to work with much lesser components, while maintaining activity. They have showcased their improved activity using not only physical but also functional readouts.

      Strengths:

      The experiments are designed in a very logical and easy to understand manner, which makes it easier not only to follow the paper, but also reproduce the results. Functional assays with the synthesized proteins are a good way to demonstrate functionality and applicability of the system. They also benchmark their system against a commercial kit to show superior performance of their system.

      Weaknesses:

      The production of the lysate requires special instrumentation, limiting accessibility.

      Comments on revised version:

      Thank you to the authors for addressing the concerns both textually and experimentally. This work has significant value.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to overcome the challenges associated with complex, conventional prokaryotic cell-free protein synthesis (CFPS) systems, which require up to thirty-five components, by developing a streamlined and efficient E. coli CFPS platform to encourage broader adoption. The main objective was to reduce the number of reaction components from thirty-five to seven, while also developing an accessible 'fast lysate' preparation protocol that eliminates time-consuming runoff and dialysis steps. The authors also sought to demonstrate the robustness and translational quality of this streamlined system by efficiently synthesising challenging functional proteins, including the cytotoxic restriction endonuclease BsaI and the self-assembling intermediate filament protein vimentin.

      Strengths:

      This study presents several key strengths of the optimised E. coli cell-free protein synthesis system in terms of its design, performance and accessibility.

      - The reaction mixture has been dramatically simplified, with the number of essential core components successfully reduced from up to thirty-five in conventional systems to just seven.

      - The "fast lysate" protocol is a significant advance in terms of procedure.

      - The system's ability to synthesise challenging, functional proteins is evidence of its robustness.

      Weaknesses:

      (1) Title: "A simplified and highly efficient cell-free protein synthesis system for prokaryotes".

      - This title is misleading since one would expect a simplified and highly efficient cell-free protein synthesis system to yield similar protein levels compared to current cell-free protein synthesis systems. What this study shows is that the composition of cell-free protein synthesis systems can be simplified while maintaining a certain level of protein synthesis. Here, optimisation does not involve maintaining protein synthesis yield while simplifying the cell-free protein synthesis system; rather, it involves developing a simplified cell-free protein synthesis system. As mentioned in my comments below, this study lacks a comparison of protein levels with a typical cell-free protein synthesis system.

      - What do the authors mean by "highly efficient"? Highly efficient compared to what experimental conditions? If one is interested by the yield of protein synthesis, is this simplified system highly efficient compared to current systems?

      (2) Figure 1, 3-5:

      - What do relative luciferase units represent? How are these units calculated?

      - In this system, the level of expression depends mainly on the level of NLuc transcripts and the efficiency of NLuc translation. How did the authors ensure that the chemical composition of the different eCFPS buffers only affected protein translation and not transcript levels? In other words, are luciferase units solely an indicator of protein synthesis efficiency, or do they also depend on transcription efficiency, which could vary depending on the experimental conditions?

      - How long were the eCFPS reactions allowed to proceed before performing the luciferase activity measurement? Depending on the reaction time, the absence or presence of certain compounds may or may not impact NLuc expression. For example, it can be assumed that tRNA does not significantly affect NLuc levels over a short period of time, and that endogenous tRNA in the lysate is present at sufficient concentrations. However, over a longer period of time, the addition of tRNA could be essential to achieve optimal NLuc levels.

      - The authors show that tRNA and amino acids are not strictly essential for the expression of NLuc, likely due to residual amounts within the cell lysate. However, are the protein levels achieved without added amino acids and tRNA sufficient for biochemical assays that require a certain amount of protein? It is important to note that the focus here is on optimising the simplicity of the buffer rather than the level of protein expression. In fact, the simplicity of the buffer is prioritised over the amount of protein produced. This should be made clear.

      - How would the NLuc level compare if all the components were optimised individually and present in an optimised buffer, compared to a buffer optimised for simplicity as described by the authors?

      (3) Line 71, Streamlining eCFPS: removal of dispensable components. This title is misleading because it creates the false impression that proteins can be produced in vitro without the addition of certain compounds. While this is true, the level of protein produced may not be sufficient for subsequent biochemical analyses. This should be made clear.

      (4) Figure 2: In the legend, change "(A) Protein expression levels of the eCFPS system measured at varying concentrations of KGlu and MgGlu2" to "(A) Protein expression levels of the eCFPS system using an Nanoluciferase (NLuc) reporter DNA measured at varying concentrations of KGlu and MgGlu2".

      (5) Lanes 302-303: "The thorough optimization of the seven core components was a critical step in achieving high protein expression levels". What are "high expression levels"? Compared to what?

      Comments on revised version.

      The authors have adequately addressed my previous concerns.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      The superiority of the optimized system might simply be due to insufficient T7 RNA polymerase in the initial lysate.

      We performed a T7 RNA polymerase titration (0–1600 ng/µL) in the initial system to test this hypothesis. Standard CFPS protocols typically utilize T7 RNA polymerase at ~90–100 ng/µL<sup>1</sup>. To fully characterize the concentration-dependent effect and determine the exact saturation threshold of T7 RNA polymerase in our system, we tested an extended range from 0 to 1600 ng/µL. As shown in the revised Figure S3B, the initial system's output reaches a plateau at ~800 ng/µL—a concentration nearly ten times higher than standard protocols. Increasing the concentration further (up to 1600 ng/µL) led to a decline in yield, likely due to inhibitory effects of excess enzyme or buffer components. Even under these T7-saturated conditions, our optimized system achieved ~45-fold higher NLuc output compared to the maximum possible output of the initial system. Notably, when the lysate concentration is increased to 70%, the productivity gap reaches nearly 80-fold, further demonstrating the extraordinary efficiency of our platform.

      As revised in the Discussion, this improvement confirms that the performance gain is not a result of a mere increase in T7 concentration. Instead, it represents a systemic synergy where our streamlined buffer and the optimized metabolic environment of the fast lysate together alleviate the transcriptional bottlenecks inherent in traditional platforms.

      Reviewer #2 (Public review):

      Performance or efficiency claims... needs to be supported by comparisons with typical cell free expression systems.

      We agree that robust benchmarking is essential for validating our claims of high efficiency. Our comparative evaluation was conducted across three levels:

      (1) Literature-based benchmarking: As detailed in Figures 3C, 4A-D, S3A-B, S4, and S5C, we extensively compared our system against the "initial" (35-component) and "PEPbased" platforms, which are established benchmarks widely utilized in CFPS literature. These diverse comparisons consistently demonstrate the superior performance and robustness of our optimized system across various conditions.

      (2) Commercial benchmarking: To provide independent verification, we performed a head-to-head comparison with a high-end commercial E. coli CFPS kit (PePExpress, Shanghai Epizyme, EC010L). As shown in the comparative data provided in this response (See author response image 1), our system exhibited remarkable rapid-expression capability, significantly outperforming the commercial kit in both speed and absolute yield. Our platform reached near-maximum yield within 2 hours, demonstrating a significant efficiency advantage over the commercial alternative.

      (3) Robustness and translational quality: The comparison was extended to challenging targets beyond standard reporters. As shown in Figures 4E-H, the successful synthesis of active BsaI restriction enzyme (a cytotoxic protein) and the functional assembly of vimentin (an aggregation-prone protein) demonstrate that our optimized system maintains superior translational quality and robustness compared to typical platforms that often struggle with such complex targets. By outperforming established academic benchmarks and a leading commercial platform in both yield and the ability to handle challenging proteins, our results provide compelling evidence that the simplified 7component system is highly efficient. In the revised Conclusion, we have explicitly contextualized "efficiency" as the integration of high protein productivity, reduced reaction complexity, and accelerated preparation speed.

      Author response image 1.

      Comparative evaluation of sfGFP yields between our _e_CFPS system (70% lysate) and a commercial kit (PePExpress) over an 8-hour time course.

      Summary of revisions: T7 titration data have been added to Supplementary Figure S3B in the revised manuscript. To provide the additional benchmarking evidence requested, commercial comparison data (PePExpress kit) are provided in Author response image 1, while the main manuscript remains focused on the mechanistic synergy and streamlined architecture of the system.

      We hope that these substantial new data and the corresponding revisions satisfy the reviewers' queries.

      References:

      (1) Kigawa, T. et al. Cell-free production and stable-isotope labeling of milligram quantities of proteins. FEBS Lett. 442, 15–19 (1999).

    1. eLife Assessment

      The study presents important findings revealing previously unresolved conformational dynamics of the heterodimeric type IV ABC transporter TmrAB using single-molecule FRET. The evidence presented is convincing, integrating careful experimental design with computational approaches to uncover states that are typically masked and difficult to detect. The work will be of interest to scientists studying the molecular mechanisms of primary active transport processes.

    2. Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

    3. Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATP-bound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single-molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states but also enabled the real-time monitoring of protein conformational changes, precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      We thank the reviewer for this accurate and thoughtful summary of our work and its broader significance. We agree that the combination of single-molecule FRET with orthogonal validation approaches enables mechanistic resolution of conformational states and transitions that are not accessible by ensemble measurements. In particular, this framework allows direct discrimination of ATP-free and ATP-bound conformations, real-time tracking of transport cycle progression, and identification of transient intermediates in the heterodimeric ABC transporter TmrAB. We further agree that these capabilities support a generalizable strategy for dissecting conformation dynamics in related ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results, and conclusions supported by the experimental data. The authors have determined the conformational dynamics of TmrAB across different ATP concentrations, including physiological ones, and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies.

      Weaknesses:

      The scientific study needs a bit of in-depth analysis with respect to consistency in K<sub>d</sub> and its implications on the mechanism.

      The apparent K<sub>d,ATP</sub> values were determined using two complementary approaches that report on different aspects of the system. Ensemble FRET measurements yielded values of 51 ± 38 µM (TmrAB<sup>NBD</sup>), 68 ± 25 µM (TmrAB<sup>PG</sup>), and 95 ± 26 µM (TmrAB<sup>PG_EQ</sup>), which are in good agreement with previously reported biochemical estimates (~100 µM for TmrAB<sup>EQ</sup>) (Stefan et al, 2020). The slightly elevated value observed for the E→Q variant may reflect modest perturbation of nucleotide handling in this slow-turnover background. Notably, the close agreement between labeled and unlabeled variants indicates that fluorophore attachment does not measurably affect ATP binding.

      In contrast, smFRET-derived K<sub>d,ATP</sub> values (13 ± 1 µM for TmrAB<sup>NBD</sup> and 2 ± 1 µM for TmrAB<sup>PG</sup>) are systematically lower. This difference likely arises from the difficulty of deconvoluting overlapping FRET populations at sub-K<sub>d,ATP</sub> concentrations, particularly for TmrAB<sup>PG</sup>, where state assignment is less well separated. Despite this quantitative offset, both approaches consistently indicate ATP saturation well below physiological concentrations and therefore support the same mechanistic conclusion that ATP binding drives conformational switching in TmrAB.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATP-bound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I have three major points and a few minor criticisms.

      We thank the reviewer for the thoughtful and constructive evaluation of our manuscript and for highlighting the strength of combining structural and single-molecule approaches. We have addressed all major and minor points in detail below and revised the manuscript where appropriate to clarify limitations, justify analysis choices, and improve transparency.

      Major points:

      (1) The main weakness is that the authors base their conclusions on a very limited set of FRET pairs. While TmrAB has been extensively studied in terms of its structure, the authors should at least acknowledge this limitation more clearly.

      We agree that our conclusions are based on a limited number of FRET reporter pairs, and we now explicitly state this limitation in the revised manuscript. The chosen labeling positions were selected to probe two functionally critical regions—the nucleotide-binding domains and the periplasmic gate—based on prior structural and spectroscopic evidence. While this represents sparse sampling of the full conformational space, it is consistent with typical smFRET studies of membrane transporters, where experimental constraints generally limit the number of simultaneously accessible labeling positions (Asher et al, 2021; Asher et al, 2022; Levring et al, 2023; Wang et al, 2020).

      Importantly, both independent reporter variants yield consistent ATP-dependent population shifts, supporting the robustness of the observed trends. We further clarify that additional labeling sites could, in principle, resolve finer structural sub-states; however, given the already limited population separation in the current variants, such extensions would likely provide diminishing returns in state resolvability under the present experimental conditions. This trade-off is now explicitly discussed.

      (2) Most smFRET distributions were fitted with one, two, or three Gaussians. However, in several cases, additional populations with noticeable amplitudes appear to be present (e.g., Figure 3c at 0.1 mM and 3 mM ATP; Figure 4a, apo; Figure 4c, 0.3 mM R9L). Could the authors clarify why these populations were not included in the analysis?

      We thank the reviewer for this careful observation. Low-amplitude sub-populations are occasionally detected in individual histograms; however, they were not included in the quantitative model because they do not meet criteria for reproducibility, amplitude robustness, or structural assignability. Specifically, these features vary between replicates, contribute minimally to total population, and cannot be mapped to structurally or biochemically defined states based on available cryo-EM (Hofmann et al, 2019), DEER/PELDOR (Barth et al, 2018; Barth et al, 2020), or accessible-volume simulations.

      Similar minor subpopulations have been reported in smFRET studies and often attributed to photophysical or labeling heterogeneity effects (Asher et al, 2022; Husada et al, 2018). To avoid over-parameterization, we therefore restricted analysis to reproducible, structurally supported states. This rationale is now clarified in the revised manuscript.

      (3) Figure 3c (3 mM ATP): Is it truly possible to distinguish the two states in this distribution?

      We agree that state separation in the TmrAB<sup>PG</sup> variant is limited (ΔE = 0.11), and we now explicitly acknowledge this constraint in the manuscript. To improve robustness under these conditions, we used a constrained fitting strategy in which the apo-state distribution was fixed from nucleotide-free measurement, reducing parameter degeneracy during fitting of ATP-bound datasets.

      While single-molecule trajectory-based approaches such as Hidden Markov Modeling would be ideal for resolving dynamic interconversion, this was not feasible due to the low fraction of dynamic traces at the available temporal resolution. We therefore rely on population-level analysis, which remains consistent across replicates and reporter variants.

      Notably, independent measurements from two reporter positions (TmrAB<sup>NBD</sup> and TmrAB<sup>PG</sup>) yield similar ATP-bound population fractions at saturating ATP concentrations (~77% vs. ~80%), supporting the robustness of the inferred state distribution despite partial overlap.

      We have revised the manuscript to more clearly articulate methodological limitations, strengthen the justification of our analytical approaches, and improve the clarity of data presentation. These revisions enhance the transparency and robustness of the study and address the reviewer’s concerns.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are a few comments that can help to improve the study.

      (1) Line 115: The authors have checked the purity and monodispersity of the protein sample using SDS-Gel and size exclusion chromatography; however, additional characterization using negative stain electron microscopy, which clearly shows the monodispersity, will be useful.

      We agree that negative stain EM can provide an additional assessment of sample homogeneity. Given the extensive prior structural characterization (Hofmann et al, 2019; Nocker et al, 2026; Nöll et al, 2017) and the SEC profiles presented here, we believe that additional negative stain EM would unlikely provide substantial new information regarding sample homogeneity. We have clarified this point in the manuscript by explicitly referencing the relevant cryo-EM studies.

      (2) Line 116: The authors have mentioned that the enzymatic activity of TmrAB was retained after purification. Although smFRET results showing conformational dynamics of TmrAB confirm its ATPase activity, a comment on the effect of labelling on ATPase activity will be useful.

      We appreciate this important point. Previous studies on spin-labeled TmrAB<sup>NBD</sup> demonstrated transport activity comparable to wild-type TmrAB, indicating that cysteine substitution and label conjugation do not substantially perturb this variant (Barth et al, 2018). In addition, AV simulations showed that fluorophores at the TmrAB<sup>NBD</sup> labeling positions do not interfere with ATP- or substrate-binding sites, supporting the conclusion that FRET labeling does not affect ATP binding, hydrolysis, or transport. For TmrAB<sup>PG</sup>, however, equivalent transport data were not available, and AV simulations suggested interference of fluorophores with periplasmic gate dynamics. We therefore directly compared the transport activity of LD555/LD655-labeled TmrAB<sup>PG</sup> and unlabeled wild-type TmrAB using a single-liposome transport assay with the fluorescein-labeled peptide C4F (RRYC<sup>F</sup>KSTEL) (<sup>F</sup>, fluorescein; Fig. 1– Fig. S3a). Both variants showed indistinguishable transport activity, demonstrating that fluorophore conjugation at the periplasmic gate preserves transport function.

      (3) Line 117 and Figure S1c. Please add the reference for consistency of ATPase activity with previous studies on TmrAB.

      We have added a reference to previous biochemical studies reporting comparable ATPase activity and kinetic parameters for TmrAB to support the consistency of our measurements.

      (4) Line 119: It mentions that "Cysteine-maleimide labeling of detergent-solubilized TmrAB achieved site-specific labeling efficiencies exceeding 90%". The legend of Figure S1d mentions about labeling efficiency in the range of 40-50%. A clarification will be helpful for the reader. Also, calculations can be extended to the ratio of LD555 and LD655 labels on the molecule, which can be considered in analyzing results.

      We apologize for the lack of clarity. The reported >90% labeling efficiency refers to the site-specific cysteine labeling efficiency per accessible site, as determined by dye incorporation. In contrast, the 40–50% values shown in Fig.1–Fig. S1d reflect the per-site efficiency for donor-lonely and acceptor-only populations respectively, which together account for the >90% overall labeling efficiency. We have revised the main text and figure legend to clearly distinguish between per-cysteine labeling efficiency and the fraction of correctly double-labeled molecules. We also clarify that only complexes with appropriate donor– acceptor stoichiometry were included in the smFRET analysis.

      (5) Figure 1: Line 627: This line mentions "For all simulations, TmrA is shown in blue with LD655 (orange) and TmrB in yellow with LD555 (green)." Is it (which label on which subunit) known for the experimental setup?

      We thank the reviewer for pointing out this potential source of confusion. In the experimental system, fluorophore attachment occurs stochastically. Therefore, the assignment of donor and acceptor dyes to specific subunits is random. The representation shown in Figure 1 reflects one possible configuration for visualization purposes only. We have clarified this explicitly in the figure legend to avoid misinterpretation.

      (6) Figure S1-2a. Tau value can be better represented in a graph for visual readers instead of in the form of a table, and a dotted line with the threshold (~1 ns) will give a better representation of no change. Values can be included in the graph as well.

      We appreciate this helpful suggestion. We have revised Figure S1-2a to include a graphical representation of fluorescence life times, including a reference line around ~1 ns to facilitate visual comparison. Numerical values are retained alongside the plot for completeness.

      (7) Figure 2a: Each component of the assembly has been pointed with an arrow, which can mix two components and confuse readers. It would be good to make a legend column on the left or right and depict or indicate each component of the assembly clearly.

      We have changed the labeling in Figure 2a to improve clarity by separating the components and introducing a clearer legend layout, ensuring that each element of the assembly is unambiguously labeled.

      (8) The physiological concentration of ATP can range up to 5-10 mM. A comment on choosing the ATP concentration specifically to be 3 mM would be useful for the readers.

      We appreciate this suggestion. While intracellular ATP concentrations can reach up to 5–10 mM, values around 3 mM are commonly used as physiologically relevant conditions in in vitro biochemical and biophysical studies. We selected 3 mM ATP as a representative near physiological concentration that ensures saturation of ATP-dependent conformational transitions while remaining comparable to previous studies on TmrAB (Hofmann et al, 2019; Nocker et al, 2026; Nöll et al, 2017; Stefan et al, 2020). We have clarified this rationale in the manuscript.

      (9) Figure 2c is not cited in the text.

      We thank the reviewer for noting this oversight. Figure 2c is now explicitly cited in the main text.

      (10) Results in Figure 2 and 3 have been analyzed using 2 and 3 Gaussian distributions, respectively. It would be good to explain the rationale for it.

      We appreciate that this important point was brought to our attention. The number of Gaussian components was determined based on the minimal model required to describe reproducible and structurally supported populations. For ATP titration experiments (Figure 2 and Figure 3), two populations (apo and ATP-bound) were sufficient and consistent across replicates. In contrast, three populations were required under trapping conditions (Figure 4), where an additional state (OFF<sup>open</sup>) becomes kinetically stabilized and clearly resolved. We have clarified this rationale in the manuscript.

      (11) Figure 3b: data points do not seem to be saturated with respect to ATP concentration. It needs more points beyond 3 mM. Different K<sub>d</sub> at different sites in the structure could represent differential local dynamics over the structure.

      Previous structural studies demonstrated that 1 mM ATP is sufficient to saturate both nucleotide-binding sites under trapping conditions (Hofmann et al, 2019), indicating that the concentration range used here is adequate. Consistent with this, both ensemble and smFRET measurements approach saturation by 3 mM ATP, a near-physiological condition commonly used in biochemical studies. While additional data points above 3 mM could further define the plateau, they are unlikely to alter the mechanistic conclusion. We have clarified this point in the manuscript.

      (12) Figure 3 and Figure 1 - S1 have two different Kd values with respect to ATP concentration; both of these graphs measure conformational changes using smFRET. A comment specifying these Kd values based on single molecule verses ensemble measurement from will be helpful for readers.

      We appreciate this important point and have clarified it in the manuscript and the response to Reviewer #1 above. The K<sub>d,ATP</sub> values in Fig. 1–Fig. S1 are derived from ensemble FRET measurements, whereas those in Fig. 3 are obtained from smFRET population analysis. This difference likely arises from the difficulty of deconvoluting overlapping FRET populations at sub-K<sub>d,ATP</sub> concentrations, particularly for TmrAB<sup>PG</sup>, where state assignment is less well separated. Despite this quantitative offset, both approaches consistently indicate ATP saturation well below physiological concentrations and therefore support the same mechanistic conclusion that ATP binding drives conformational switching in TmrAB. We now explicitly distinguish these methods and their interpretation in the manuscript.

      (13) Figure 4: Slow-turnover TmrAB mutant has been employed in cysteine mutant on the PG opening side, but not towards the NBD side. Either experimental data or a comment on not pursuing it would be helpful for the reader. Similarly, experiments in the presence of peptide and in the absence of ATP, which can help to understand the role of substrate in conformational dynamics in the absence of ATP, are not pursued in this study. Along similar lines, experiments with wild type, in the presence of MgADP +/- substrate, are not shown in this study.

      We thank the reviewer for these insightful suggestions. The slow-turnover variant was specifically applied to the periplasmic gate reporter (TmrAB<sup>PG</sup>) because this construct provides direct sensitivity to outward-facing conformations, which are central to resolving the OF<sup>open</sup> state. In contrast, the NBD reporter primarily monitors nucleotide-binding domain (NBD) dimerization and is less suitable for distinguishing periplasmic conformational differences.

      Experiments in the absence of ATP but in the presence of peptide, as well as MgADP ± substrate, would indeed be valuable for further dissecting substrate effects. However, these conditions are beyond the scope of the current study, which focuses on ATP-driven conformational dynamics and the identification of kinetically hidden intermediates. We have added a statement in the Discussion to acknowledge these possibilities as directions for future work.

      (14) Figure 4, peptide concentration has been varied in the right panel. The result can also be presented as the % of OFopen and OFoccluded state with increasing concentration of peptide.

      We thank the reviewer for this suggestion. While such a plot would indeed be informative and could improve our understanding of substrate binding and substrate-induced trans-inhibition, the current dataset does not contain sufficient data points to construct a reliable concentration-dependent curve, particularly given that peptide saturation was not reached in our experiments. The characterization of substrate binding is further complicated by the presence of two distinct substrate-binding sites one in the outward-facing and one in the inward-facing state with likely completely different K<sub>d</sub> values and would require a more complex binding model. We have therefore decided against including this plot in the current manuscript. We do acknowledge, however, that future smFRET studies with improved temporal resolution are particularly well suited to investigating substrate binding to TmrAB and its effects on conformational equilibrium, and we have noted this in the Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) In all figures, can you please label the transporter schematics with the conformational states they represent?

      We thank the reviewer for this suggestion. All transporter schematics in the main and supplementary figures have been updated to include clear labels indicating the corresponding conformational states, thereby improving clarity and consistency.

      (2) As a suggestion, it may improve clarity to include the labelling positions (residue numbers) directly in Figure 1a and b, even though they are provided in the legend.

      We appreciate this suggestion. Residue numbers corresponding to labeling positions have now been added directly to Figure 1a and b to improve readability and facilitate interpretation.

      (3) Lines 183-188: This is a key point. It would be helpful to include a reference line for the expected state (0.63). Interestingly, this value coincides with the shoulder observed in Fig. 3c (0.1 mM ATP). Is there an explanation for this (see also point 2)?

      We thank the reviewer for highlighting this point. We considered adding a reference line at 0.63 to the plot; however, we decided against it. While a subpopulation does appear at ~0.63 —consistent with the expected FRET efficiency of the OF<sup>open</sup> conformation—it is only present in a single condition (0.1 mM ATP) and is not observed across other ATP concentrations for this TmrAB variant. It more likely reflects a minor non-reproducible subpopulation or photophysical artefact, in line with our response to Point 2 of the public review (Reviewer #2).

      (4) The final section of the Results section seems like an afterthought, especially since the heading suggests a broader scope.

      We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      References

      Asher WB, Geggier P, Holsey MD, Gilmore GT, Pa; AK, Meszaros J, Terry DS, Mathiasen S, Kaliszewski MJ, McCauley MD, Govindaraju A, Zhou Z, Harikumar KG, Jaqaman K, Miller LJ, Smith AW, Blanchard SC, Javitch JA (2021) Single-molecule FRET imaging of GPCR dimers in living cells. Nat Methods 18: 397–405. doi:10.1038/s41592-021-01081-y

      Asher WB, Terry DS, Gregorio GGA, Kahsai AW, Borgia A, Xie B, Modak A, Zhu Y, Jang W, Govindaraju A, Huang LY, Inoue A, Lambert NA, Gurevich VV, Shi L, Lefkowitz RJ, Blanchard SC, Javitch JA (2022) GPCR-mediated beta-arrestin activation deconvoluted with single-molecule precision. Cell 185: 1661– 1675 e1616. doi:10.1016/j.cell.2022.03.042

      Barth K, Hank S, Spindler PE, Prisner TF, Tampé R, Joseph B (2018) Conformational coupling and transinhibition in the human antigen transporter ortholog TmrAB resolved with dipolar EPR spectroscopy. J Am Chem Soc 140: 4527–4533. doi:10.1021/jacs.7b12409

      Barth K, Rudolph M, Diederichs T, Prisner TF, Tampé R, Joseph B (2020) Thermodynamic basis for conformational coupling in an ATP-binding cassette exporter. J Phys Chem LeJ 11: 7946–7953. doi:10.1021/acs.jpclett.0c01876

      Hofmann S, Januliene D, Mehdipour AR, Thomas C, Stefan E, Brüchert S, Kuhn BT, Geertsma ER, Hummer G, Tampé R, Moeller A (2019) Conformation space of a heterodimeric ABC exporter under turnover conditions. Nature 571: 580–583. doi:10.1038/s41586-019-1391-0

      Husada F, Bountra K, Tassis K, de Boer M, Romano M, Rebuffat S, Beis K, Cordes T (2018) Conformational dynamics of the ABC transporter McjD seen by single-molecule FRET. EMBO J 37: e100056. doi:10.15252/embj.2018100056

      Levring J, Terry DS, Kilic Z, Fitzgerald G, Blanchard SC, Chen J (2023) CFTR function, pathology and pharmacology at single-molecule resolution. Nature 616: 606–614. doi:10.1038/s41586-023-05854-7

      Nocker C, Pečak M, Nocker T, Fahim A, Sušac L, Tampé R (2026) Single-molecule dynamics reveal ATP binding alone powers substrate translocation by an ABC transporter. Nat Commun 17 doi:10.1038/s41467-026-70021-1

      Nöll A, Thomas C, Herbring V, Zollmann T, Barth K, Mehdipour AR, Tomasiak TM, Bruchert S, Joseph B, Abele R, Olieric V, Wang M, Diederichs K, Hummer G, Stroud RM, Pos KM, Tampé R (2017) Crystal structure and mechanistic basis of a functional homolog of the antigen transporter TAP. Proc Natl Acad Sci U S A 114: E438–E447. doi:10.1073/pnas.1620009114

      Stefan E, Hofmann S, Tampé R (2020) A single power stroke by ATP binding drives substrate translocation in a heterodimeric ABC transporter. eLife 9: e55943. doi:10.7554/eLife.55943

      Wang L, Johnson ZL, Wasserman MR, Levring J, Chen J, Liu S (2020) Characterization of the kinetic cycle of an ABC transporter by single-molecule and cryo-EM analyses. eLife 9: e56451. doi:10.7554/eLife.56451

    1. eLife Assessment

      The authors present important evidence for a WIPI2-Retriever complex (termed CROP2) that couples cargo selection to carrier fission at endosomes. CROP2 appears to function analogously to the previously described CROP1 complex, formed by WIPI1 and Retromer, with which it shares structural similarities. They provide compelling evidence that CROP1 and CROP2 regulate the trafficking of distinct subsets of cargoes; however, the cellular evidence for the existence of these distinct complexes is mostly inferred from immunoprecipitation analysis and would benefit from further validation.

    2. Reviewer #1 (Public review):

      WIPI1 is a PROPPIN family protein that has been implicated in Retromer-mediated membrane fission events. Although the cargos that it has been tested to be important for are diverse, one of the cargos that is unaffected is Beta1-Integrin. This leads the authors to assess another PROPPIN family protein - WIPI2, which is a homolog of WIPI1. KD using siRNA is effective and had no consequences on LAMP1, EGFR trafficking or GLUT1 trafficking. Integrin-B1, however, had a large and significant defect in its recycling from the endosome, with a clear endosomal colocalisation. Complementation experiments with WT WIPI2 recovered the phenotype, but various mutant WIPI2 complements resulted in elongated tubules, and there was also a dominant negative effect of the mutant. Integrin is a classic retriever cargo, so the authors rationalise that WIPI2 may be playing a role with retriever that WIPI1 plays with retromer. To assess this, they perform a set of immunoprecipitations. SNX17, the retriever-associated sorting nexin, co-IPs with WIPI2 in a VPS26C-dependent manner. VPS26C but not VPS26 co-IPs with WIPI2, and the reciprocal with WIPI1. These interactions were not present for the FSSS mutation of WIPI2. WIPI2 localises to Rab11 endosomes mainly, as does retriever. Mutations of WIPI2 not only affected WIPI2 localisation, but also VPS35L mutations, indicating that there is a functional relationship between the two.

      Comments on revised version.

      The reviewers have responded appropriately to all the points. I have no remaining concerns.

    3. Reviewer #3 (Public review):

      Summary:

      The manuscript of Mayer and colleagues analyzes the function of WIPI proteins in mammalian cells. The authors identified previously CROP as a complex consisting of WIPI1 and the retromer complex, primarily in yeast cells. In mammalian cells, both WIPI1 and WIPI2 exist, whereas retromer has a homologous complex termed retriever. The now find that WIPI2 can form a complex with retriever subunits. They name this complex CROP2. Their data further indicate that CROP2 and CROP1 have distinct substrate specificities as knock down of CROP2 subunits affect beta1 integrin sorting, whereas knock down of CROP1 affects EGFR and GLUT1. The further identify a similar sequence (FSSS) in both WIPI1 and WIPI2, which is required for their specific binding to retromer and retriever.

      Strengths:

      CROP1 and CROP2 seem to use similar features for their formation, and have different substrates, which is convincingly shown.

      Weaknesses:

      The analysis lacks information that this is a complex as claimed. It can be deduced from the immunoprecipitation analysis.

      Comments on revised version.

      The authors answered my questions and adjusted the text accordingly. Figure 10 was not part of the submitted version. It should be checked by the editor.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      WIPI1 is a PROPPIN family protein that has been implicated in Retromer-mediated membrane fission events. Although the cargos that it has been tested to be important for are diverse, one of the cargos that is unaffected is Beta1-Integrin. This leads the authors to assess another PROPPIN family protein - WIPI2, which is a homolog of WIPI1. KD using siRNA is effective and had no consequences on LAMP1, EGFR trafficking or GLUT1 trafficking. Integrin-B1, however, had a large and significant defect in its recycling from the endosome, with a clear endosomal colocalisation. Complementation experiments with WT WIPI2 recovered the phenotype, but various mutant WIPI2 complements resulted in elongated tubules, and there was also a dominant negative effect of the mutant. Integrin is a classic retreiver cargo, so the authors rationalise that WIPI2 may be playing a role with retreiver that WIPI1 plays with retromer. To assess this, they perform a set of immunoprecipitations. SNX17, the retreiver-associated sorting nexin, co-IPs with WIPI2 in a VPS26C-dependent manner. VPS26C but not VPS26 co-IPs with WIPI2, and the reciprocal with WIPI1. These interactions were not present for the FSSS mutation of WIPI2. WIPI2 localises to Rab11 endosomes mainly, as does retriever. Mutations of WIPI2 not only affected WIPI2 localisation, but also VPS35L mutations, indicating that there is a functional relationship between the two.

      On the whole, I find the manuscript compelling. The manuscript is very clearly written, the results are convincing and well performed. The flow of experiments is logical, and although not comprehensive in the subsequent mechanistic understanding, the fundamental findings are important and convincing. My comments below are, on the whole, minor and are intended to support the communication of the findings to the field.

      We are happy that the reviewer has received our work quite positively.

      (1) The IP interaction data were convincing; however, for me and some others, an interaction is only convincing when performed in vitro, and understood at a structural level. I do not suggest the authors do that in this case; however, I think, at a minimum, some sensible moderation of claims would be useful here.

      Indeed, quantitative in vitro data on the affinities would be a nice addition. However, we have significant trouble to recombinantly express and purify well-behaved WIPI2 in sufficient quantities for such studies. We keep working in this direction but are not there yet.

      We have now inserted a phrase into the discussion section highlighting this limitation: "Our immunoprecipitation assays cannot distinguish and more detailed structural and interaction studies with pure compounds will be necessary to elucidate the nature of this interaction". We nevertheless think that the the isoform specificity of the IPs, the effect of the point mutations in WIPI2 on these interactions, and the functional effects in vivo lend signficant support to the notion of a complex even if there is no proof of direct binding of WIPI2 to Retriever.

      (2) I found the final localisation data and its interpretation confusing. My interpretation of that data would not be that the retreiver is relocalised, but rather that there is less of both recruited to the membrane and the remaining localisation distribution is shifted. In addition, I am not quite sure of the model here - is the idea that WIPI2 recruits retreiver, if that is the case, I find it hard to resolve with its role as a mediator of fission. Clarity would be appreciated here.

      We are not quite sure what "final" localisation data the reviewer refers to, but we guess it is Fig. 9. This figure primarily provides in vivo evidence supporting the connection between Retriever and WIPI2. It does this by showing that the S67 substitution shifts both proteins. In WIPI2 wildtype cells, WIPI2 and VPS35L strongly colocalize in Rab11 compartments. S67 substitutions in WIPI2 abolish this localisation; WIPI2 shifts mainly to Rab5 compartments, where VPS35L shows only a moderate increase, and to Rab7 compartments, where VPS35L shows no increase at all.

      We do not understand the reviewer's interpretation that less Retriever would be recruited to the membranes in the S67 variants. VPS35L remains completely associated with punctate, presumably membrane-bounded structures also in the mutants, providing no evidence for a detachment from the membrane. The same is observed in a WIPI2 knockdown. Therefore, we did not claim that WIPI2 is the main factor recruiting Retriever to the membrane, for which our experiments yield no hints. This does not exclude that the interaction of WIPI2 could strengthen membrane recruitment, or that two pools of Retriever exist, one interacting with Snx17 and another interacting with WIPI2, and that both link to each other in a coat. We did not dwell on this in the discussion because our experiments cannot distinguish these possibilities and were not conceived to analyse membrane recruitment of Retriever.

      (3) I am concerned that the repeats being compared for statistical analysis are not biological repeats but technical repeats (cells in the same experiment). I should think the idea of the statistical comparison is to show experimental reproducibility and variability across biological repeats. Therefore, I would expect an appropriate number of biological repeats (3 or more minimum), to be the data compared in the statistical analysis and graphs. I think it is appropriate to average the technical repeats from each biological repeat. I find these to be useful resources https://doi.org/10.1083/jcb.202401074, https://doi.org/10.1083/jcb.200611141

      The repeats being compared are biological repeats from independent experiments. This is described in Methods, where the reviewer may not have seen it. In order to make the independent experiments more evident in the figures, we have now colour coded the individual cell measurements from the three independent experiments. This allows to visualize both the individual data points, the average from each experiment and the variability across the independent experiments.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from De Leo and Mayer presents evidence that the PROPPIN protein, WIPI2, associates with the Retriever complex, and is required for the proper transport of the SNX17-Retriever cargo, beta1-integrin. This finding fits with prior papers from the Mayer lab, which showed that a related PROPPIN, WIPI1, is required for the transport of some SNX27-Retromer cargo, including GLUT1. The retromer and retriever complexes are architecturally similar. Importantly, they act at the same endosomes, and each transports cargo from endosomes to the plasma membrane. Thus, the possibility that each also requires a structurally related PROPPIN is of interest. However, the manuscript is incomplete, and the main claims are only partially supported.

      Strengths:

      The topic that PROPPIN proteins are important for the function of the Retromer and Retriever complexes expands our view of the trafficking complex.

      Weaknesses:

      Many important controls are missing. Several points that are made in the manuscript are only supported through a single approach.

      We made a serious effort and implemented many suggestions of this reviewer, but orthogonal approaches are not always available or accessible.

      Reviewer #3 (Public review):

      Summary:

      The manuscript of Mayer and colleagues analyzes the function of WIPI proteins in mammalian cells. The authors previously identified CROP as a complex consisting of WIPI1 and the retromer complex, primarily in yeast cells. In mammalian cells, both WIPI1 and WIPI2 exist, whereas retromer has a homologous complex termed retriever. They now find that WIPI2 can form a complex with retriever subunits. They named this complex CROP2. Their data further indicate that CROP2 and CROP1 have distinct substrate specificities as knockdown of CROP2 subunits affects beta1 integrin sorting, whereas knockdown of CROP1 affects EGFR and GLUT1. They further identify a similar sequence (FSSS) in both WIPI1 and WIPI2, which is required for their specific binding to retromer and retriever.

      Strengths:

      CROP1 and CROP2 seem to use similar features for their formation, and have different substrates, which is convincingly shown.

      Weaknesses:

      The analysis lacks information that this is a complex as claimed. It can be deduced from the interaction analysis, but was not shown.

      It is of course desirable to obtain a detailed structural and in vitro characterisation of this interaction, which we have not provided because we currently do not have sufficient amounts of well-behaved source material for this. We nevertheless think that the interaction we show, which is strictly isoform-specific and dependent on single amino acid substitutions in a motif that in CROP1 is necessary for the interaction its recombinant subunits, supports that CROP2 is a similar a complex. We don't show a direct interaction but also don't claim in the manuscript that the interaction between WIPI2 and Retriever is direct and independent of additional factors.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, the reviewers generally value the contribution to the field, but they feel that some claims require additional experimental support.

      (1) I have summarized the major points below.

      (a) Both reviewers 1 and 2 agree that the quality of localization data presented in Figure 9 and S5-S7, and the interpretation of the data, could be improved. See comment 2 from reviewer 1 and comments 23, 24 and 25 from reviewer 2. They not only suggest ways to improve the presentation of the data, but additionally suggest improving the staining of the Rab11 marker and additionally explain the lack of co-localization between VPS35 and Rab5, which has been reported in the literature.

      This impression was due to the fact that some figures showed projections of image stacks, which was not indicated clearly in the figure legend. We have changed this and now show single image planes throughout all figures.

      (b) Both reviewers 1 and 3 note that the evidence supporting a functional WIPI2-Retriever complex in vivo is currently weak. We agree that additional biochemical data demonstrating the presence of the CROP1 and CROP2 complexes in vivo would strengthen the central message of the paper and elevate it to a more fundamental discovery.

      We understood that the reviewers did not ask for further in vivo evidence but would welcome structural characterisation of the complex and quantitative binding data in vitro with purified proteins. Structural characterisation is out of scope of our study and in vitro binding studies have remained hampered by the fact that WIPI2 is hard to express and purify and not well behaved in vitro.

      (c) All reviewers agree that the authors should carefully repeat their statistical analysis to account for the number of biological replicates. Reviewer 1 suggests publications that the authors could refer to.

      The reviewers have probably overlooked the respective description in the methods section, where it had been stated that we analysed biological replicates from independent experiments. In graphs showing measurements from individual cells we now make this evident through colour coded dots, in which each colour represents data points stemming from an independent experiment. This makes it evident that the variance from experiment to experiment is low. The means (n = 3) were generally compared using a two-tailed unpaired t-test.

      (d) Reviewer 2 additionally has various minor points that would greatly improve the readability and presentation of the work, and we recommend addressing (comments 1, 2, 3, 4, 12, 15, 17, 20, 27, 28, 29). All reviewers, in general, provide great minor suggestions. It would be great if the CROP1 and 2 complexes could be clearly introduced in each figure. We also agree that the WIPI2 CT labelling is confused and should be changed to "control" or similar.

      Many of the points raised by this reviewer were actually quite minor or questions of personal preference, not major problems as stated in the review. Nevertheless, we found a number of useful suggestions in this review and have addressed these points as detailed in the response to reviewer 2.

      (2) In addition to the major shared concerns laid out in the points above, reviewer 2 has some further minor suggestions:

      (a) Comment 6. Could the author explain the discrepancies between the example blot shown in Figure 1D and the quantification (1E).

      The two have actually been quite consistent. The reviewer might have mistaken the marker lane as the 0 min reference value to arrive at this impression. We have now removed the marker lane to avoid this.

      (b) Comment 9 - could the authors clarify how surface labelling experiments were carried out?

      This had been clearly described in the methods section, where this reviewer has probably not seen it.

      (c) Comment 11 - The reviewer suggests normalizing the surface levels of markers to the cell area and not per cell. This is a reasonable suggestion.

      The analysis had already been performed as proposed. This had been clearly described in the methods section, which the reviewer may not have looked at.

      (d) Comment 19 "In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures." The authors could try a Rab4 or Rab11 overexpression plasmid to show whether these are elongated recycling tubules.

      This has now been added.

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The figures are not colourblind friendly, and should be changed to be so. Additionally, single colour images should be grayscale.

      That was a good learning opportunity. We adapted the colour schemes of the images to make them more colourblind friendly, now using magenta, green, and white for the overlaps. In doing so we have relied on published recommendations, but we have not found a colourblind colleague to check the efficacy of this change.

      (2) WIPI2^CT labels are confusing, as people may think they are a mutant. I suggest changing to "control" or similar.

      These have been changed.

      (3) "The effect was comparable to that of a knockdown of SNX17 (Figure 3 A, B)." On page 6. Based on this sentence, I was expecting to see a comparison to SNX17 KD, but it was not there as far as I can tell.

      This statement referred to a publication by P.Cullen and collaborators. We have changed the wording and inserted the (missing) reference to make this clear.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is modest. In addition, many of the claims should be better supported by the addition of orthogonal data. Moreover, the quality of some of the data presented needs to be improved. Overall, the manuscript requires better descriptions of the methods. In many figures, it was not clear how the experiments were performed.

      The experimental descriptions that the reviewer refers to had been provided in the Methods section, where this reviewer may have overlooked them.

      The paper should also be better organized. Some less important findings are in the main figures, whereas some critical results are in the supplemental figures. In addition, there were multiple issues with the readability of the paper, and the authors should consider using a professional editor to make the paper easier to read.

      We had given the paper to colleagues who found it clear, and also Reviewer 1 has underlined its clarity. Nevertheless, we have re-phrased the manuscript in some parts to optimise it.

      One of the main claims in the paper is that the FSSS motif of WIPI2, as well as a conserved amphipathic helix, is critical for WIPI2 function in the CROP2 complex. It is notable that these are the same regions that are also critical for the role of WIPI2 in autophagy (Gubas et al., 2024 PMID: 39152217). The authors should include this information in the manuscript and cite the paper.

      Indeed. We mention this now in the introduction of the revised version.

      Additional Major Issues:

      While some of the issues raised below are actually minor and/or matters of personal preference, several comments led us to improve and correct the figures and we thank this reviewer for the constructive suggestions.

      (1) In Figure 1, it appears from the representative images that WIPI2 KD cells have higher levels of EGFR (Figure 1A and 1B). Is this correct?

      To some degree. This increase is not systematic. A moderate increase has been observed only in 2 experiments out of 4. Therefore, we did not investigate this.

      (2) Also in Figure 1, the colocalization is difficult to see. The authors should add the separate channels in addition to the merged images. Since the point is supposed to be that there is no impact on EGFR, all of this data could go into the supplement.

      We had considered this already for the original version but dismissed the idea. The overlap is quantified in Fig. 1C, which provides the relevant values from four experiments. Fig. 1A/B provide only sample pictures, which also permit to see overlap (yellow) 0 and 5 min after the induction of degradation, which vanishes at later timepoints. Separating the channels would quadruple the space that this figure occupies, which would not be practical and not change the point to be made.

      (3) The scale bars for each panel differ from each other. To better assess the data, the exact same magnification should be shown for each panel.

      Corrected

      (4) Figure 1C is confusing. The authors should explain which lines correspond to EEA1 and LAMP1.

      Corrected

      (5) In Figure 1D, the authors show different blots for control and WIPI2 KD. Could the authors compare WIPI2 and EGFR in the same blot? Without a comparison on the same blot, it is impossible to know whether the starting levels of EGFR are the same. Moreover, the quantitation in Figure 1E sets the value for each cell line to 100%. Instead, the starting levels in each cell line should be compared. The authors should use the amount of EGFR at zero time in the control cells to define 100%, and then indicate the relative initial EGFR levels in the WIPI2KD cells.

      A new blot is shown now and the quantification has been performed as proposed.

      (6) The quantification in Figure 1E does not match the representative blot shown in Figure 1D. According to the graph, the rate of degradation of EGFR is similar in both cell lines. But the representative blot shows that there are large differences.

      We do not understand this comment. The representative blot shows similar kinetics for both. Perhaps the reviewer got confused by the fact that a marker lane was still present on the left blot and not labelled as such. The new version of the figure corrects this.

      (7) The blot showing the WIP2 knockdown in Figure 1D has a lot of background. However, the blot of the WIPI2 knockdown in Figure S1 looks very good. The authors should make sure that they load enough sample and use a good antibody for the experiments in Figure 1.

      The new blot that we added in response to comment 5 corrects this.

      (8) In Figure 2 and Figure 3A, the cells are too confluent. This is an issue because the cells might not be metabolically active. In addition, the signal is saturated. The authors should make sure that all of the data is collected on cells that are not too confluent.

      The confluency of the culture cannot be judged from single frames, which were selected to show several cells. We had controlled confluency and underlined in the Methods section that “For microscopy, the cells were plated on 18-mm-diameter glass coverslips on 24-well plates and grown for 2 or 3 days according to the protocol of DNA or siRNA transfection by reaching a confluency of 70-80%”. The reviewer may not have seen this.

      (9) One main issue with these figures, especially the non-permeablized cells, is that it is impossible to assess how much of the signal is on the cell surface. The authors should provide the methods that they used to prevent inadvertent permeabilization of the cells. Were these experiments performed at 4 degrees? The authors should include a control of an antibody to a protein that is not found on the cell surface.

      There is an internal control in that the non-permeabilised WIPI2KD cells, which have been treated with the same antibody, show no much less staining than the control cells (Fig. 3A). In WIPI2KD cells, integrin becomes accessible for antibody staining only upon detergent permeabilization. This demonstrates that our procedure does not lead to significant inadvertent permeabilization of the cells.

      (10) The authors should perform surface biotinylation assays as an orthogonal approach to determine GLUT1 levels and beta1-integrin levels at the cell surface, respectively.

      There is a strong, qualitative difference in the surface labelling of beta1-integrin that is not observed for GLUT1. Given that, it is not obvious to us what additional argument would be provided by surface biotinylation or subfractionation experiments.

      (11) In quantifying surface levels of GLUT1 or beta1-integrin by microscopy, the authors should normalize to the cell area, rather than per cell.

      The reviewer has probably not seen that the Methods section states that the cell area has been used for normalisation.

      (12) In Figure 3, the nuclear DAPI stain in the KD cells is much less bright than in the control cells. The authors should make sure to choose representative images.

      The nuclear DAPI signal has been visible in all cells. Depending on the position of the nucleus, is shape and dimension in the z-direction, individual nuclei can show different degrees of staining. The images shown are representative. We have adjusted the settings now to make the nuclei in the WIPI2KD cells easier to spot.

      (13) For the immunofluorescence studies, the authors should be using single z planes rather than maximum projection.

      Images have been exchanged by single planes.

      (14) For the experiments in Figure 3, the authors should check the total levels of EEA1 and LAMP1 by western blot to test whether WIPI2 KD affects the levels of these proteins. If these organelle marker proteins are impacted, this could impact the colocalization measurements shown in Figures 3C and D.

      We have measured the total fluorescence intensity of EEA1 and LAMP1 in the images. It shows no significant difference between control and WIPI2 knockdown cells (new Fig. 3F, H).

      (15) In Figure 4A, the helical representation is rotated in the WIPI2-Sloop; the orientation of the residues that are not mutated should stay the same.

      Yes. Done.

      (16) In Figure 4B and 4C, cells that were not transfected with WIPI2 WT or WIPI2 Sloop should be shown.

      Since the transfection efficiency is limited, the fields contain both non-transfected (lacking green fluorescence) and transfected cells (showing green fluorescence). We have now marked transfected cells with an asterisk.

      (17) The cells in the lower panel of 4B have an unusual morphology and are much more round. The authors should choose cells that are representative of each experimental condition.

      We now provide another field.

      (18) In Figure 4C, it looks like the magnification of the top panels is different from the bottom panels. The same magnification for all the panels should be shown (and the size of the scale bars should be the same.

      Corrected

      (19) In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures.

      We have done this (new Fig. S4). Rab4 is on tubules. Rab5 on the structures from which the tubules emanate.

      (20) In Figure 5A, the top scale bar is missing.

      Corrected.

      (21) In Figure 5B, the confluency is too high.

      See our response above. A single field does not permit to judge this. Confluency was controlled for all cultures. The cultures were not confluent.

      (22) The IP studies shown in Figures 6, 7 and 8, should be accompanied by colocalization studies.

      Colocalization measurments have now been integrated into the manuscript (Figs. S5, S6). They are consistent with the IP data.

      (23) Figure 9 was very confusing and should be broken up into multiple figures. Data showing that localization did not change in any of the cell lines can be put in figures that are distinct from figures that show that localization changed in the various mutants. Figures that show no change can go in the supplement.

      Since every panel of Fig. 9 shows a statistically significant difference we left the figure unchanged.

      (23) Representative figures should be shown in the same figure as the corresponding graph. In addition, the order of the colocalization data shown in the graphs and figures should match the order described in the text.

      We consider the graphs of Fig. 9 as the relevant information. Representative images are just illustration. Integrating them with the graphs would make it necessary to split everything up into multiple figures, making it harder to compare the different combinations. Therefore, we left the figures unchanged.

      (24) In Figure S7, the Rab11 signal looks continuous, which makes the colocalization analysis meaningless. The authors should determine how to take images that can be evaluated. On a more minor note, the zoomed panels should be labeled as well.

      This is a result of having shown a projections of multiple planes. The images have now been replaced by single plane images. Zoomed panels have been labelled and the scale bar added.

      (25) The low colocalization of VPS35L with Rab5 is surprising, as SNX17 has been previously shown to co-localize with early endosomes positive for EEA1. This result may have occurred due to overexpression because the authors chose to utilize plasmids that express a tagged protein. There are antibodies to each of the endogenous proteins, and this is what should be used for this set of experiments.

      This comment made us control the analysis performed for these images, which by mistake had been performed on z-projections rather than on single planes. This distorted the values. The re-analysed data shows a higher colocalisation with Rab5, but it remains inferior to colocalisation with Rab11.

      (26) The authors should determine whether β1-integrin colocalizes with WIPI2 in endosomal compartments.

      This was done. WIPI2 colocalizes with beta-integrin on EEA1-and SNX17-positive strcutures but not positive for LAMP1 (Fig. 3E/F).

      Minor points

      (27) In one of the panels in Figure 1A, "30 min" is duplicated.

      Removed

      (28) In Figures 5C and 5D, the y-axis should indicate that this is surface β1integrin.

      Changed and added “surface”

      (29) In Figure 9 there is a typo in panel A. It is VPS35L and not VPS35.

      Corrected

      Reviewer #3 (Recommendations for the authors):

      This is an overall convincing study, which shows that the two complexes, CROP1 and CROP2 function at different membranes and serve different substrates. While I agree with their localization analysis, I have one key issue. The authors claim that each of the two forms a complex and base this on their specific pull-down and western blot analyses.

      I find it important that they show that both indeed form stable complexes in vivo, using pull-down and mass spectrometry approaches. They have all the necessary tools in hand and could use WIPI1 and WIPI2 to demonstrate the existence of the two complexes. The FSSS mutants of each are good controls for such an analysis.

      The manuscript actually presents the demanded in vivo experiments. Figs. 6 to 8 show pull-downs of WIPI1 and WIPI2 from cells, including also the FSSS mutant. While we haven't analysed this interaction by mass spectrometry, the Western blot analysis confirms the analysis. Cooperation of these proteins is further supported by the in vivo phenotypes, where the S67A substitution in WIPI2 produces a similar phenotype on integrin beta1 localisation as inactivation of Retriever.

      A second aspect is the general presentation. The paper would be a lot more accessible if the subunits of each complex (CROP1 and CROP2) were also introduced in the figures of each part. For readers, a final model is helpful to put the data into context and show where each complex operates in the cell.

      We have introduced a scheme of the respective complexes, including the names of the compunds, in Figs. 6 and 7 to avoid confusion.

      Finally, it is not clear how the statistics compare to repeats in their data. This should be clarified.

      This had been described in methods. Statistics has always been done on biological replicates stemming from independent experiments. We have added a cartoon (Fig. 10) depicting the trafficking pathways affected by CROP1 and CROP2.

    1. eLife Assessment

      This paper reports the findings of a neuroimaging experiment that tested the hypothesis that the cortex, specifically early visual areas, reinstates certain content from past episodic events. This is a useful study that highlights the role of early sensory cortices in supporting rapid, one-shot learning of location information for long-term memory. The strength of the evidence is solid, with the methods, data, and analyses broadly supporting the claims.

    2. Reviewer #1 (Public review):

      Summary:

      This paper reports the findings of a neuroimaging experiment that tested the hypothesis that the cortex, specifically early visual areas, reinstates the content from single events during our lives. The researchers tested this hypothesis by presenting to-be-remembered pictures of objects at spatial locations on the computer screen and then testing subjects with both recall and recognition. They show that during memory testing, the spatial location of the object can be decoded from the pattern of cortical BOLD responses measured with fMRI. They go on to show that the spatial tuning is higher during recognition than recall, that the tuning is correlated with memory retrieval accuracy, and that the retrieved precision is predicted by the encoded precision, particularly in the higher-level visual areas. Thus, the paper finds evidence of cortical reinstatement of details from a single event in a human life.

      Strengths:

      This is a strong manuscript that I have had the luxury of commenting on during a round of review at another prestigious journal. As a result, the authors have already made changes to address previous comments about highlighting the complementary learning systems approach more to motivate the alternative prediction that the cortex should only show evidence of reinstatement after repeated presentations. In addition, the authors have fleshed out the discussion of working memory in this task. They also revised their review of the literature to include citations suggesting spatial locations are normal parts of our episodic representations, likely obligatory in nature, as my group and others have argued in completely unrelated work. I applaud the authors for being responsive to a previous round of review and using the comments to address relatively minor issues with the paper, even though they moved on to a different journal. Thus, I found the paper even stronger than at first approach, and at first blush, the results were intriguing and the paper well written.

      Weaknesses:

      There is a logical perspective in the narrative that seems to unnecessarily weaken the paper. The paper shows evidence consistent with the conclusion that mnemonic representations are contained in early visual cortex, but then argues that those representations are not actually stored therein. For example, the first half of the last sentence of the conclusions (see page 19 of the manuscript). I understand the perspective that subcortical mechanisms must be involved in the act of retrieval, given the neuropsychology and other evidence. But if storage is elsewhere with the same fidelity so as to code this information, then how would such a memory system work? The MTL neurons would need to have the real, precise representation of all the orientations encoded at all the retinotopic locations, a mirror to V1 in terms of precision, because that's the actual memory representation being retrieved, so its fidelity will be limited by what is stored in the file, so to speak. Then, at retrieval, the paper proposes that the brain just reactivates the encoding context in V1 to help with the response output and ensure the precision of the behavioral responses. This must mean that the hippocampus/MTL has cells and networks with tuning functions that match the precision in all the cortical sensory systems that they are integrating context across, given the episodic memory models like Polyn and colleagues (2009, Psych Rev). So, there are little MTL maps that are completely redundant with V1, M1, A1, S1, etc.? Why such redundancy?

      Why not propose that what the subcortical systems do is to encode a unique pattern for that episode, that is separated from others, that just links (or provides pointers to, in computer science jargon) the contextual details stored in the cortical networks themselves? In this way, we can explain why neglected patients also neglect their memories of the town square. This has always been my interpretation of the results of the Polyn et al. (2006, Science) paper and the models tested with those whole-brain results. That is, you see widespread cortical context reinstatement during (one-shot) free recall events that included visual selective cortex for faces when faces were being recalled, but included a broad network, probably V1, and activating sounds in A1, body posture in M1, etc., though the latter three examples did not discriminate between categories of memoranda, in their experiments. Given that you show that activity in V1 during retrieval looks like it is being used, you should propose that the early cortex really participates in memory storage functions. V1 neurons are wired up to neurons of other selectivities in a competitive network with plastic synaptic connections. How would experience be prevented from changing activity in the cortex? Yes, cortical changes slow after the critical periods, as studied in the classic eye suturing experiments to study ocular dominance, but changes in cortical representations do not stop with maturity, with the pinwheel centers looking like they are context sensitive, thus, changing rapidly to events across time (Okamoto, Ikezoe, et al., 2011, Sci Reports). The brain would need a no-plasticity mechanism, and instead, it looks like the cortex can completely rewire even in adulthood (Buonomano & Merzenich, 1998, Annu Rev Neuro).

      I believe that the paper needs to describe the strong/radical interpretation of the current findings; that they are consistent with the view that the entire brain may be a memory structure, with encoding linking representations across sensory cortices. But also activating semantic and lexical systems, emotional networks encoding those aspects of context which we know can sometimes strongly drive effects, a nice prediction that could be made in the discussion/conclusions. Here you are looking at how precise the visual reinstatement is in V1 during retrieval following one exposure. One parsimonious mechanism to explain this effect is that the brain stores details of events using the neurons that do the high-fidelity perception of the event. Given that our goal is to stimulate thinking among fellow scientists so that this paper can be a citation classic, I think the paper should be revised so that it paints a complete picture of the theoretical possibilities of its findings.

    3. Reviewer #2 (Public review):

      Summary:

      The study aims to show that the early visual cortex is not merely a sensory-perceptual region that encodes stimuli while they are physically present, but also supports the formation and retrieval of long-term episodic memories. Instead, the authors demonstrate that spatially tuned reactivation of early visual cortex after a single encoding event supports memory-guided behavior, such as recalling an object's original location.

      Strengths:

      The study provides solid evidence that location information for single, trial-unique objects is reinstated in early visual cortex during both recognition and recall, even without explicit spatial demands, and the remembered vs. forgotten analyses link spatial tuning to behavior. The one-shot design and absence of explicit spatial instructions are important strengths that bring the paradigm closer to everyday, incidental episodic experiences and go beyond highly trained cue-target associations.

      Weaknesses:

      (1) Conceptually, the main findings would appear less surprising without a sharper theoretical contrast. Given basic retinotopic coding, it is natural that object identity and location are jointly encoded when an object is presented at a particular position, so spatially tuned reinstatement in V1-V3 can be interpreted as a reconfirmation of known properties unless more clearly contrasted with theories that emphasize more abstract, position-invariant cortical representations following hippocampal-cortical recoding. As currently framed, the introduction does not fully articulate what existing accounts might predict, or what pattern of results would have challenged those accounts, which somewhat weakens the perceived theoretical payoff.

      (2) It also remains somewhat unclear why early visual cortex (V1-V3), specifically, is the critical locus for the spatial information of interest, as opposed to higher-level visual or parietal regions that could also provide a spatial scaffold; clearer rationale and, if possible, control analyses in additional regions would help here.

      (3) Since gaze behavior is central to any spatial account, it would be helpful to report basic eye-tracking analyses comparing remembered versus forgotten trials, especially at encoding, to rule out systematic differences in fixation patterns that could contribute to the spatial tuning results.

    4. Reviewer #3 (Public review):

      Summary and Overall Evaluation:

      This is an elegant paper addressing an important question: whether spatial location is automatically activated during the recall of object memories. Building on prior work that relied on trained or repeated stimuli, the present study uses unique objects with one-time encoding across four spatial locations - a meaningful advance in ecological validity. The experimental design is clean, the data analysis is well-executed, and the reported effects, while small, are intriguing and open up interesting questions about the role of spatial structure in visual memory. Overall, this is a solid contribution, and my comments below are intended to help the authors strengthen the paper further.

      Major Comments

      (1) Incidental encoding.<br /> Was the memory task fully incidental - that is, were participants unaware that a subsequent memory test would follow encoding? This seems important for interpreting the automaticity claim that is central to the paper's contribution, and should be clarified explicitly.

      (2) Spatial extent of the analysis - higher visual regions and negative pRFs.<br /> The analysis appears restricted to regions V1-V3. Have the authors examined higher visual areas as well? This seems like an important omission given that object memory likely engages regions well beyond the early visual cortex. Relatedly, recent work by Adam Steel and colleagues suggests that spatially tuned negative pRFs may play an important role in memory. Have the authors considered examining these? Expanding the analysis in these directions could substantially enrich the findings.

      (3) Mechanism - retinotopic or spatiotopic?<br /> The paper makes a compelling case that spatial structure supports memory, but the nature of that spatial structure deserves more discussion. Are the effects retinotopic or spatiotopic in nature? The current design may not be able to fully dissociate these possibilities, but this distinction is theoretically important, and the authors should engage with it directly. Even a careful discussion of what the current data can and cannot tell us on this point would be valuable.

      (4) Relationship between encoding failure and retrieval failure.<br /> For trials where memory performance is worse, and the encoding models fail, is there a systematic relationship between how the pRFs fail at object retrieval versus spatial retrieval? In other words, are the pRFs wrongly tuned in the same way at both stages? This analysis could provide meaningful insight into whether object and location retrieval draw on shared spatial representations.

      (5) Object shape and spatial mapping.<br /> Real-world objects vary considerably in surface structure and shape, which may affect how cleanly they map onto a specific spatial location. Was this considered in the analysis? What was taken as the correct or peak location for each object, and how was this defined when objects extended across space? Apologies if this was addressed in the methods and I missed it.

      (6) Time course of pRF activation.<br /> Is there a way to examine the time course of pRF activation within a trial? Do the spatially tuned responses arise immediately upon retrieval, or do they build up over time? Even a preliminary analysis of this would be of considerable theoretical interest, as it would speak to whether spatial reinstatement is an early automatic process or a later, more deliberate one.

      (7) Effect size and functional significance.<br /> The authors acknowledge that the reported effects are very small, which I appreciate. However, this does raise genuine questions about functional significance that I think deserve a more direct response. One approach that would help contextualize the spatial effects would be to compare their magnitude to that of another feature - object identity, for example - to give readers a sense of the relative importance of spatial versus non-spatial information in memory representations. I recognize this may not be straightforward with the current design, but even a brief discussion of how one might benchmark the spatial effects would be helpful.

      (8) The attention account.<br /> I found the discussion of attention less than fully convincing. The authors appear to argue against an attentional interpretation of the spatial effects, but it is not clear why participants wouldn't attend to the encoded location during retrieval - particularly in a design with relatively few retrieval cues, where spatial location may be one of the most useful available. The attention account thus seems difficult to rule out on the basis of the current data, and the discussion should engage more seriously with this alternative rather than setting it aside.

      (9) Later-remembered versus later-forgotten objects - BOLD signal.<br /> Were later-remembered objects associated with stronger overall BOLD responses during encoding compared to later-forgotten objects, or was the effect specific to the pRF modelling? Clarifying this would help readers understand whether the spatial effects are part of a broader pattern of stronger encoding or something more specific to the spatial reinstatement mechanism.

    1. eLife Assessment

      This fundamental study provides convincing evidence that distinct molecular mechanisms underlie AAV-associated retinal toxicity in retinal pigment epithelial cells and photoreceptors, advancing our understanding of gene therapy-related retinal injury. The authors employ a rigorous and comprehensive experimental approach, including multiple knockout mouse models, transcriptomic analyses, and genetic loss-of-function studies, which substantially strengthen the mechanistic conclusions. Some concerns remain regarding vector characterization, the absence of procedural injection controls, and the limited interpretation of adult versus neonatal studies; nevertheless, the study makes a substantial contribution to the field and provides a strong foundation for future translational investigations.

    2. Reviewer #1 (Public review):

      This study examines the mechanisms underlying retinal toxicity associated with certain AAV gene therapy vectors, particularly in the retinal pigment epithelium (RPE) and photoreceptors following expression of transgenes such as GFP. The findings suggest that AAV-related retinal toxicity is driven less by transgene identity itself and more by distinct pathogenic mechanisms, including stress-induced injury in RPE cells and interferon-mediated damage in photoreceptors. The comments are as follows:

      (1) The AAV vectors were manufactured in-house, and the production method is described in sufficient detail. However, were any characterization assays performed beyond qPCR-based titer determination, such as vector genome titer, capsid titer, empty/full capsid ratio, sterility, bioburden, endotoxin, mycoplasma, residual host cell DNA, residual plasmid DNA, or residual host cell protein testing? These analyses, particularly those assessing residual impurities and microbial contamination, are critical, as such contaminants may provoke inflammatory responses following subretinal injection. This, in turn, could confound the interpretation of the results, including the identification of the molecular pathways contributing to toxicity as well as the specific role of GFP-associated toxicity. Please provide any characterization information for the AAV vectors.

      (2) The study uses contralateral or uninjected eyes as controls, but this choice may not adequately account for changes induced by the subretinal injection procedure itself. Because the earliest assessment of RPE toxicity was performed at 2 weeks post-injection, any injury, inflammation, retinal detachment-related stress, or wound-healing responses triggered by the surgical procedure could have contributed to the observed phenotype. As a result, comparisons to uninjected eyes alone make it difficult to distinguish vector or transgene-specific toxicity from procedure related effects. Inclusion of a more appropriate procedural control, such as sham-injected eyes or eyes injected with vehicle/buffer alone, would have strengthened the study by enabling clearer discrimination between injection-related retinal responses and toxicity attributable to the AAV construct or transgene expression.

      (3) The authors used phalloidin staining on RPE-choroid flatmounts to evaluate RPE toxicity, which provides useful information on RPE morphology and structural disruption. However, it would be highly informative to also assess the presence and distribution of subretinal microglia/macrophages, for example, by Iba1 immunostaining, in the same preparations. Specifically, determining whether Iba1-positive cells accumulate in or around areas of RPE dystrophy would help clarify the contribution of local inflammatory responses to the observed pathology. Such analysis could strengthen the interpretation of the toxicity phenotype by revealing whether RPE degeneration is accompanied by focal immune cell recruitment and whether these cells spatially associate with regions of tissue damage. This would also provide additional insight into whether inflammation is likely to be a downstream consequence of RPE injury or a more direct contributor to disease progression, especially in light of publications by Danial Saban's group regarding the characterization of microglia phenotypes using RNA-seq analysis.

      (4) The Discussion should also address the anatomical and procedural differences between neonatal and adult mouse eyes, particularly with respect to retinal thickness and the potential impact of subretinal injection-related injury. Because the RPE toxic effects appeared less severe in adult mice, it would be valuable for the authors to consider whether this difference reflects true age-dependent biological susceptibility or, at least in part, differences in the mechanical consequences of the injection procedure. Neonatal retinas are thinner and structurally less mature than adult retinas, which may render them more vulnerable to injection-associated stress, retinal detachment, or secondary tissue injury following subretinal delivery. In contrast, the greater retinal thickness and maturity of the adult eye may provide some degree of resilience to procedural trauma, thereby reducing the apparent severity of RPE damage. Expanding the Discussion to consider these factors would strengthen the interpretation of the age-related differences observed in toxicity and help distinguish vector- or transgene-driven effects from potential confounding effects introduced by the delivery method itself.

      Overall, this manuscript presents a detailed and comprehensive analysis of transgene-induced retinal toxicity and makes effective use of multiple mouse models to dissect the contribution of relevant molecular pathways. The study is particularly strengthened by its systematic approach, combining histologic, transcriptomic, and genetic loss-of-function strategies to distinguish the mechanisms underlying toxicity in the RPE versus photoreceptors. By evaluating several knockout mouse lines, the authors can move beyond descriptive observations and begin to assign causality to specific stress and immune signaling pathways, thereby providing important mechanistic insight into AAV-associated retinal injury. These findings are timely and relevant to the broader field of ocular gene therapy, as they highlight the complexity of vector- and transgene-related toxicity and underscore the need for careful pathway-level evaluation during preclinical development.

    3. Reviewer #2 (Public review):

      Summary:

      Adeno-associated viruses (AAVs) are popular gene therapy vectors, but AAVs can cause toxicity. This is particularly evident following expression of some transgenes, e.g., GFP, in the retinal pigment epithelium (RPE), which leads to loss of RPE cells and photoreceptors. Here, we sought to unravel the toxicity mechanism(s). Several transgenes, self and non-self, were tested for toxicity, with no clear correlation for this variable. RPE RNA-sequencing revealed upregulation of translational processes, cell stress, cytokine release, antiviral responses, and leukocyte infiltration pathways. Toxicity-inducing pathways were explored for causality by injecting toxic AAVs into mice deficient for intrinsic, innate, or adaptive immune pathways. The CHOP KO partially alleviated toxicity for RPE but not photoreceptors, whereas the type I interferon receptor KO partially alleviated toxicity for photoreceptors but not RPE. In situ hybridization of interferon pathway transcripts (IFNB1, IFNAR1) revealed that the RPE and retina can produce and potentially respond to interferon. These data suggest that transgene-induced cell stress responses in the RPE lead to RPE cell death, while interferon signaling contributes to the death of photoreceptors.

      Strengths:

      This manuscript used numerous KO mouse models to evaluate the interferon pathway, inflammatory cytokine pathways, the complement pathway, toll-like receptor signaling, cytosolic DNA sensing, double-stranded RNA sensing strain, intrinsic cellular stress pathways, as well as strains deficient for B cells and T cells or B cells, T cells, and natural killer cells. This is a robust piece of work with rigorous controls, groups, and timepoints tested. The RNA-sequencing data provided helpful guidance on the pathways that should be assessed when analyzing AAV toxicity to the retina.

      Weaknesses:

      The main weakness of the study is that it focuses on subretinal administration to neonatal mice, and the canonical TLR9-MyD88 was not found to have an impact on the AAV toxicity measured. More information could have been provided to understand the discrepancy.

    1. eLife Assessment

      This study presents a useful methodological advance that better enables the simultaneous measurement of gene expression and chromatin accessibility in individual cells. The evidence supporting the improved detection of gene expression is solid. The method has the potential to be more broadly impactful if it were expanded to include orthogonal validation strategies. This method will be of interest to those studying transcription and gene regulation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. In the latest version, the authors have made textual revisions that note caveats about the quality of the chromatin accessibility data.]

      In the manuscript entitled "Flexible and high-throughput simultaneous profiling of gene expression and chromatin accessibility in single cells," Soltys and colleagues present easySHARE-seq, a method described as an improvement upon SHARE-seq for the simultaneous measurement of RNA transcripts and chromatin accessibility.

      The authors demonstrate the utility of easySHARE-seq by profiling approximately 20,000 nuclei from the murine liver, successfully annotating cell types and linking cis-regulatory elements to target genes. The authors claim that easySHARE-seq supports longer read lengths potentially enabling better variant discovery or allele-specific signal assessment, though they do not provide direct evidence to support these specific claims.

      A key strength of the protocol is enhanced sequencing efficiency, achieved by shortening the Index 1 read from 99 to 17 nucleotides. This reduction does not come at a significant cost to barcode diversity, retaining approximately 3.5 million combinations. Additionally, the approach allows for the sequencing of a sub-library to assess quality prior to final barcoding and sequencing which seems quite clever.

      While the increase in RNA transcript recovery is substantial, it appears to come at a cost: there is a notable decrease in ATAC fragments per cell compared to the original SHARE-seq (and other platforms). Likely as a result, the dimensionality reduction (UMAP) shows good resolution for RNA profiles but relatively poor resolution for accessibility profiles. Furthermore, the presented data suggests potential ambient RNA contamination; specifically, the detection of Albumin in HSCs and B cells is likely an artifact of the protocol rather than a biological signal.

      Overall, the study is well-presented and represents a promising advance.

    3. Reviewer #2 (Public review):

      Aims:

      The authors sought to optimize SHARE-seq, a multimodal single-cell method, to improve the simultaneous profiling of gene expression and chromatin accessibility. Their goal was to enhance barcode design for better sequencing efficiency and cost savings, while improving overall data quality. They then applied their optimized method, easySHARE-seq, to study liver sinusoidal endothelial cells (LSECs) to demonstrate its utility in examining gene regulation and spatial zonation.

      Strengths:

      The improved barcode design is an advance, increasing the proportion of sequencing reads dedicated to biological information rather than barcode identification. This modification offers practical benefits in terms of sequencing costs and read length, potentially reducing alignment errors. The method also demonstrates improved RNA detection compared to the original SHARE-seq protocol. The biological applications showcase how simultaneous measurement of both modalities enables analyses that would be practically impossible with single-modality approaches, particularly in examining how chromatin states change along developmental or spatial trajectories.

      Weaknesses:

      There is a notable reduction in chromatin accessibility detection compared to the original SHARE-seq method, likely limiting the use of the method in certain situations.

      Overall:

      The authors achieve their aim of creating an optimized protocol with improved barcode design and enhanced RNA detection. The method represents a useful advance for specific experimental contexts where the trade-offs are appropriate.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Comments from Reviewing Editor:

      I want to share that both reviewers appreciated that this revision has appropriately addressed many of the concerns they raised. However, reviewers concurred that additional wet-lab experiments which validated the findings would have made the work much more impactful; and their concerns about the quality of chromatin accessibility data appear not to be fully resolved. Might I suggest a textual revision that specifically points out these caveats, if you are not able to provide additional data? This would then proceed to VOR without additional need to review. Thanks much for your patience while I assessed the manuscript claims and reviewer opinions.

      The changes were very minor (2 sentences in the Discussion and a small section in the Supplementary Notes). It would be great if we could proceed to the VOR stage.

    1. eLife Assessment

      This important study shows that NPAS4, a gene that is switched on by neural activity, enhances the spatial and temporal precision of hippocampal neurons during navigation. These findings, based on selective and sparse gene deletion, are supported by convincing evidence. However, the experiments were performed entirely in animals exposed to long-term environmental enrichment, which leaves open the question of whether the same effects would emerge under standard housing conditions. This study will be of interest to neuroscientists studying neuronal circuits and spatial coding.

    2. Reviewer #1 (Public review):

      Summary:

      NPAS4 is an activity-dependent transcription factor that regulates inhibitory synapses onto active pyramidal neurons. In this study, the authors examined whether this molecular mechanism influences neural coding in awake animals. To accomplish this, they generated a sparse, CA1-specific NPAS4 knockout in mice and compared knockout neurons with neighboring wild-type neurons recorded from the same animals during navigation. They found that, although neurons lacking NPAS4, which received diminished somatic inhibition and enhanced dendritic inhibition, still encoded location, their spatial firing was less precise: place fields were broader and less stable, showed weaker firing within the field, and exhibited more firing outside the field. KO neurons also exhibited poorer temporal organization with weaker coupling to theta oscillations and reduced phase precession, two signatures of precise spike timing in the hippocampus. Overall, the study suggests that NPAS4 links the balance of somatic and dendritic inhibition to the quality of circuit-level coding by refining the spatial and temporal precision of neuronal firing.

      Strengths:

      Using a sparse CA1-specific knockout, the authors compared NPAS4-deficient neurons with neighboring wild-type neurons within the same animal and network. This is a significant advantage because it minimizes confounding factors arising from global circuit disruption, providing a clearer comparison of genotypes. Furthermore, the rigorous optogenetic tagging strategy used to distinguish KO from WT neurons in vivo makes the single-cell comparisons much more convincing. Electrophysiological recordings from intermingled WT and KO neurons enable precise spike-timing measurements relative to a shared local field potential, which would be challenging to obtain with calcium imaging.

      Weaknesses:

      Rather than an acute manipulation, the authors rely on a chronic, sparse knockout, and NPAS4 had been deleted for at least one month before recording. Consequently, while the paper demonstrates a robust long-term phenotype, it is less definitive about the immediate causal sequence by which NPAS4 induction alters inhibition and reshapes spatial and temporal coding. Furthermore, the study focuses on single-neuron coding during navigation and does not test whether the observed degradation in coding precision leads to corresponding impairments in learning or memory in the same animals. In the discussion, the authors suggest that NPAS4 may be especially important for ripple-associated activity during sleep; however, the paper does not test this possibility.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Payne and colleagues examines how cell-autonomous loss of the activity-dependent transcription factor NPAS4 reshapes spatial and temporal coding in CA1 pyramidal neurons of behaving mice. The work builds on the Bloodgood lab's established framework in which NPAS4 reorganizes inhibition along the somatodendritic axis of CA1 pyramidal cells, principally by regulating CCK+ basket cell synapses, and asks whether this transcriptionally driven reconfiguration of inhibition propagates into the spike-train statistics that underlie hippocampal function. The combination of sparse Cre delivery with channelrhodopsin-mediated optotagging in Npas4 fl/fl:Ai32 mice is technically elegant, as it permits within-animal comparisons of intermingled wild-type and knockout pyramidal neurons sharing a common LFP, which is a significant analytical advantage for spike-timing analyses and for controlling network-level confounds. The reported phenotype is internally consistent and converges on a coherent story: knockout neurons exhibit broader and less stable place fields, lower signal-to-noise within fields, increased out-of-field activity, weaker theta-phase coupling, and shallower phase precession slopes, with the temporal deficits at least partly explained by enlargement of the spatial receptive field.

      Strengths:

      Several aspects of the work deserve explicit recognition. The validation of the optotagging strategy is thorough, including the high-power stimulation control to corroborate WT classification and the post hoc histological alignment of GFP+ density with electrophysiologically identified KO fractions. The decision to test NPAS4 function in adult mice maintained in long-term enriched environments addresses an important gap, since most prior work has focused on juveniles or short-term induction paradigms. The acute slice recordings recapitulating the somatodendritic inhibition phenotype reassure the reader that the in vivo measurements are interpreted against a known synaptic substrate. The analytical framework, especially the difference maps across epochs and the linear regression decomposition of phase precession slope into genotype, field size, and theta modulation strength, is rigorous and goes beyond simple group-level comparisons. The conceptual contribution, namely the demonstration that an activity-dependent transcription factor can be tied to single-neuron coding properties in vivo, is meaningful, although it is fair to note that the direction of the effect, given that the CCK to place cell link and the NPAS4 to CCK link have each been established in prior independent studies, is largely along the lines one would predict.

      Weaknesses:

      The most consequential concern, in my view, is the experimental context in which the entire study is conducted. Every animal is housed in an enriched environment for two to three months, and Figure 1A itself shows that NPAS4 expression in CA1 is essentially undetectable in standard-environment conditions and only emerges with enrichment. This raises the question of whether the manuscript is in fact describing the function of NPAS4 in general, or the function of NPAS4 specifically as recruited by chronic enrichment. The paper, in its current framing, elides this distinction and presents the EE state as if it were the baseline, which it is not. EE is known to alter hippocampal connectivity, the dynamics of place cell ensembles, and the expression of many activity-dependent genes; the CCK to pyramidal cell connectivity that the authors invoke as the mechanistic anchor is also dense in standard housing, so the absence of detectable NPAS4 in SE conditions raises the further conceptual problem of how NPAS4-negative neurons would normally be innervated by CCK+ basket cells in the first place. A direct comparison of WT and KO neurons in standard-environment animals, even on a smaller scale, would discriminate between two very different interpretations, namely that NPAS4 has a constitutive role in tuning CA1 firing versus that it is specifically engaged by enrichment-driven activity and contributes to an EE-specific reorganization of coding. Recent work, including Chiaruttini and colleagues (2025), reports baseline NPAS4 expression in CA1, so the SE result in Figure 1A may itself underestimate normal expression and deserves further scrutiny. Without an SE comparison, the generality of the conclusions cannot be assessed, and the title and abstract risk overstating the scope of the findings, particularly when one considers that NPAS4 is also induced by contextual fear conditioning and other paradigms, which would predict context-specific effects rather than a uniform refinement function.

      A closely related concern is the meaning of the knockout itself. Even under EE, only a few percent of CA1 pyramidal neurons express detectable NPAS4 at any given moment (Figure 1A), yet the AAV strategy deletes the gene in 30 to 60 percent of pyramidal neurons. In effect, the majority of cells classified as KO in this study would not have been expressing the protein under the relevant conditions, so the population that is statistically driving the WT versus KO differences must include a non-trivial fraction of neurons in which the deletion has no protein-level consequence. This dilutes the expected effect and raises a more interesting biological question: are the observed phenotypes carried by the few KO neurons that would have expressed NPAS4, or do they emerge from a constitutive function of the gene that is broader than the IHC signal suggests? An additional, related possibility is that NPAS4 expression segregates non-uniformly across functional classes, for example, concentrating in cells with particular firing-rate or spatial-tuning profiles, in which case the "KO" label is binary at the level of the manipulation but graded at the level of biological consequence. Stratifying the KO population by some proxy of activity history, or relating the magnitude of the phenotype to per-cell measures of recent firing, would help address this. As written, the manuscript treats the KO designation as homogeneous, while the underlying biology is almost certainly not.

      A third concern, more conventionally statistical, is the treatment of cells as independent observations. The analyses rely almost uniformly on Kolmogorov-Smirnov tests applied to individual units pooled across animals, but cells recorded in the same animal share not only a common subject but a common network, since WT and KO neurons here are intermingled in the same CA1 microcircuit. Cell numbers per animal range widely, so a mixed-effects framework treating animal as a random factor, or a hierarchical bootstrap, would clarify which effects are robust against animal-level and session-level variability and protect against pseudo-replication. This concern is particularly acute for the smaller effects in Figure 2C-E, where the cumulative distributions overlap substantially, and the differences could plausibly be driven by a small number of mice or sessions. In several figures, the individual dots in supplementary panels are not labeled by animal or session, and that information would be useful for assessing how much of each effect is carried by which subset of the cohort.

      The absence of a Cre/ChR2 expression control is a separate gap. The comparison throughout the manuscript pits Cre+ ChR2+ neurons (NPAS4 KO) against neighboring non-transduced neurons (WT). This is internally elegant, but leaves open the possibility that part of the phenotype arises from chronic ChR2 expression or constitutive Cre activity rather than from NPAS4 loss, especially given that most of the readouts are subtle. A small companion cohort of Ai32 mice without the floxed Npas4 allele, injected with the same AAV and processed through identical optotagging and electrophysiology pipelines, would address this definitively and is, in my view, a near-essential addition.

      Several of the downstream phenotypes would benefit from stratified comparisons that hold first-order properties constant. Many of the downstream differences (stability across epochs, theta coupling, phase precession) could, in principle, be inherited from the upstream difference in firing rate, since the high-firing and high-spatial-information cells in the WT pool are likely contributing disproportionately to the group statistics. The authors do perform firing-rate-matched controls in Figure S4D-G, which is helpful, but the analysis should be extended in two ways: a parallel stratification by spatial information for the stability analyses in Figure 4, and matched comparisons of theta coupling (Figure 5) and phase precession (Figure 6) on neurons drawn from overlapping firing-rate and spatial-information distributions. The regression decomposition for phase precession is a step in this direction and shows that field size, not genotype, is the dominant predictor of slope; this finding, in my reading, deserves more prominent framing in the discussion than it currently receives, since it implies that the temporal precision phenotype is largely downstream of the spatial one rather than a parallel deficit.

      The place field stability analysis is interesting but somewhat under-analyzed. The authors show that KO fields shift toward the field entrance more rapidly than WT fields and propose that this reflects an accelerated or dysregulated Mehta-effect-like dynamic. The framing is attractive, but the analysis does not establish that the shifts are systematic in the same way the classical Mehta effect is. An alternative reading is that the elevated out-of-field firing creates spurious local maxima that the peak-finding procedure occasionally classifies as field shifts, especially when in-field firing is reduced. A control analysis using a fixed reference window around the original peak, rather than re-identifying the peak each epoch, would help distinguish a genuine plasticity-like shift from instability driven by noise. The behavior of the WT population in epoch 4 also raises a question: would the drift intensify over longer recording windows, and to what extent is the apparent drift imposed by the repetitive structure of the task itself, in which animals are effectively running on a constrained linear /circular track that may impose drift-like dynamics across the population independently of genotype?

      A final note on mechanism. The manuscript leans on prior work showing that NPAS4 regulates CCK+ basket cell synapses, and uses this as the mechanistic anchor for the coding deficits. The connection is reasonable but remains indirect within this study, since the authors do not measure CCK+ interneuron activity, perisomatic inhibition, or local circuit dynamics in the same animals. The discussion already acknowledges some of this, but the speculative framing of dendritic versus somatic inhibition contributions could be tightened, especially given that competing inhibitory sources (PV+ basket cells, axo-axonic cells, OLM interneurons) also shape the spatial and temporal features measured here. A more cautious mechanistic framing, distinguishing what is demonstrated from what is inferred from prior work, would be appropriate.

      In summary, this is an ambitious and technically demanding study that makes a meaningful contribution by linking activity-dependent transcriptional regulation of inhibition to the spatial and temporal organization of CA1 spike trains in awake, behaving mice. The within-animal optotagging design is a real strength, the phenotype is internally consistent across multiple coding metrics, and the conceptual implications for how experience tunes single-neuron coding are significant. The principal concerns, namely the unaddressed enrichment confound that pervades the entire dataset, the conceptual ambiguity around what a KO designation actually means at the cell level when only a small fraction of CA1 neurons express the protein, the statistical treatment of nested observations from a shared microcircuit, the missing transgene control, the absence of stratified comparisons by firing rate and spatial information for the secondary phenotypes, and the somewhat overreaching mechanistic framing of the discussion, are all addressable, and if handled carefully would substantially strengthen the manuscript. With these revisions, the work would be a valuable contribution to the literature on how the molecular memory of activity shapes circuit-level coding.

    4. Author response:

      We appreciate the time and attention to our manuscript and the feedback from the reviewers, who were overall supportive of the work. Both reviewers validated the technical approach we used to differentiate the wild-type (WT) and knockout (KO) neurons noting: “The combination of sparse Cre delivery with channel rhodopsin-mediated optotagging in Npas4 fl/fl:Ai32 mice is technically elegant” and “the rigorous optogenetic tagging strategy used to distinguish KO from WT neurons in vivo makes the single-cell comparisons much more convincing.” Furthermore, they note the consistency of the reported results, stating: “The reported phenotype is internally consistent and converges on a coherent story”.

      Both reviewers also pointed out several concerns or points of improvement for the manuscript. Below, we first offer several scientific and methodological clarifications that we believe resolve a number of the reviewers' concerns. We then outline which remaining points we plan to address through revision, and which fall outside the scope of the current study.

      Scientific Clarifications:

      Request for a standard housing control. Both of the reviewers brought up the long-term enrichment paradigm (EE) we opted to use for this study and expressed interest in seeing data from standard housed (SE) animals. This is an approach the lab has taken in its slice physiology work [1-3], where comparing EE and SE conditions has revealed important differences between cellular phenotype. However, the in vivo experiments described here differ in a key way: obtaining these recordings requires extensive handling, training, and daily transport between the vivarium, home cage, and behavior room. These experimental steps themselves constitute the kind of novel, salient experience known to induce NPAS4, making a true SE comparison unattainable within this paradigm. In our experiment, mice were housed in EE as a supplemental, well-established strategy to induce NPAS4 in CA1 pyramidal neurons but we believe the behavior alone would be sufficient. We will describe this more clearly in the text of the manuscript.

      Consistent with this view, place fields recorded from wild-type mice in other studies using SE but undergoing comparable handling and training procedures, are similar in size, spatial information, and stability to the WT place fields we reported here [4,5]. As part of our revisions, we will consider statistical comparisons between our WT neurons and those reported in other studies to quantitatively assess whether a difference exists.

      More broadly, we note that the existing literature on NPAS4 induction does not, to our knowledge, establish a baseline level of NPAS4 expression in CA1 pyramidal neurons in the complete absence of behavioral experience. Reports of NPAS4 expression in CA1 have generally relied on animals exposed to some form of salient or novel experience [3,6,7], consistent with our framework that NPAS4 induction reflects behaviorally-driven activity rather than a constitutive baseline.

      Expression profile of NPAS4. Reviewer #2 brought up a concern about the extent of the NPAS4 expression, referring to the IHC results in Figure 1A stating: “Even under EE, only a few percent of CA1 pyramidal neurons express detectable NPAS4 at any given moment (Figure 1A), yet the AAV strategy deletes the gene in 30 to 60 percent of pyramidal neurons. In effect, the majority of cells classified as KO in this study would not have been expressing the protein under the relevant conditions.” We wish to clarify two points here. First, in the experimental paradigm used to obtain the IHC results, mice were exposed to enrichment for only 90 minutes while in the in vivo physiology paradigm, mice were housed in an enriched environment (with frequent toy changes to ensure novelty) for weeks. Thus, NPAS4 is almost certainly expressed in a much larger percentage of WT neurons in mice that were kept in chronic enrichment and used for the in vivo studies. Second, while the NPAS4 protein is only expressed in cells for several hours following neuronal activity, it initiates an inhibitory synapse phenotype that persists long-term. Thus, even though a small percentage of neurons are NPAS4+ in the IHC results, it is likely that a much larger percentage of them have expressed NPAS4 in the past and now show the inhibitory synapse phenotype. Evidence for this comes from the slice physiology results in Figure 1C (and see similar results from adolescents [1-3]) in which animals were housed in enrichment long-term and differences between inhibition persisted in nearly every WT/KO comparison.

      We also recognize the related possibility that NPAS4 expression may not be uniform across the pyramidal cell population, but may instead concentrate in particular functional subtypes, such as cells with higher firing rates or stronger spatial tuning. As part of our revisions, we plan to test this directly by stratifying the KO population by firing rate and relating it to the magnitude of the observed phenotype. Taken together, we believe that while only a small fraction of CA1 pyramidal neurons are NPAS4+ at any given moment, a much larger fraction have experienced NPAS4 induction and the accompanying synaptic reorganization over the timescale of chronic enrichment making the WT/KO comparison in this study substantially less diluted than the IHC snapshot alone would suggest.

      Timeline of NPAS4 expression and synaptic reorganization. Reviewer #1 pointed out that this study only examines the effects of NPAS4-deletion on longer timescales (weeks to months after the virus expression and subsequent knockout) stating “[the study] is less definitive about the immediate causal sequence by which NPAS4 induction alters inhibition and reshapes spatial and temporal coding”. The reviewer is correct, the temporal relationship between NPAS4 expression, changes in synaptic inhibition, and changes in neuronal firing are important outstanding questions in the field. Currently, we lack molecular tools that would enable us to clearly test these relationships but with our existing, albeit limited information, we have the following working model.

      When an animal is placed into a new context, a subset of CA1 pyramidal neurons will fire action potentials in a spatially refined manner. This activity will drive NPAS4 expression in those neurons, resulting in protein expression that persists for a couple of hours before the protein is degraded.

      Following expression, NPAS4 will bind to various sites in the genome and initiate a genetic program which results in changes in inhibition recruiting CCK basket cell synapses to the soma and destabilizing CCK dendritic synapses. The exact mechanism behind this reorganization of inhibition is unknown, but the phenotype likely emerges over the course of several hours following NPAS4 expression and persists for days following the stimulus that induced NPAS4.

      While our chronic knockout approach does not allow us to resolve the precise timing of events in this sequence, it does allow us to ask a distinct and complementary question: what is the long-term consequence for a neuron that has never been able to execute this program? Our results demonstrate that NPAS4-deficient neurons which cannot initiate NPAS4-dependent inhibitory reorganization regardless of their activity history show systematic degradation in spatial and temporal coding precision. This establishes that the NPAS4-dependent inhibitory phenotype has lasting and functionally meaningful consequences for in vivo information encoding, a question that shorter-timescale or acute manipulations would not be well-positioned to address. Resolving the immediate causal sequence between NPAS4 induction, synaptic reorganization, and changes in firing will be an important goal for future work as new molecular tools become available.

      Behaviors that drive NPAS4 expression. Reviewer #2 pointed out that “NPAS4 is also induced by contextual fear conditioning and other paradigms which would predict context-specific effects rather than a uniform refinement function.” They are correct NPAS4 is expressed in response to different behavioral paradigms, including fear conditioning and environmental enrichment. However, the subregion in which NPAS4 is induced depends critically on the behavioral paradigm. When mice are exposed to contextual fear conditioning, NPAS4 expression is robust in CA3 and the dentate gyrus but negligible in CA1 [6]. This is consistent with the known activity patterns of these subregions: CA3 neurons are strongly recruited during contextually-dependent associative learning, while CA1 neurons are more reliably driven by exposure to novelty and respond in a spatially-refined manner. Consistent with this, studies using fear conditioning have focused on behavioral discrimination and synaptic changes in CA3 and granule cells [6]. To our knowledge no study has examined the relationship between fear conditioning, NPAS4, and CA1 pyramidal neuron function. Whether behavioral paradigms beyond environmental enrichment and spatial navigation can induce NPAS4 in CA1, and what consequences that might have for pyramidal neuron firing, are interesting questions for future work.

      We also wish to address the conceptual framing underlying this concern. In CA1, we do not believe that “context-specific effects” are separable from a “uniform refinement function.” CA1 pyramidal neurons respond in a context-dependent manner. When a mouse is placed onto a linear track, there is a subset of neurons that will increase their activity over the course of that exposure. But within this subset, individual neurons will also show spatially-refined responses firing action potentials as the animal runs through the corresponding place field. The spatial precision NPAS4 confers is always nested within context-dependent mechanisms NPAS4 refines whatever representation a neuron is already computing, rather than overriding the context-dependency of that representation. We therefore do not view these as competing frameworks.

      The role of NPAS4 in shaping CCK synapses. Reviewer #2 made the point that “the CCK to pyramidal cell connectivity that the authors invoke as the mechanistic anchor is also dense in standard housing, so the absence of detectable NPAS4 in SE conditions raises the further conceptual problem of how NPAS4-negative neurons would normally be innervated by CCK+ basket cells in the first place.” We wish to clarify that NPAS4 is not necessary for the formation of CCK synapses onto CA1 pyramidal neurons there are likely a number of NPAS4-independent mechanisms that regulate this synaptic connectivity (for example, see [8]). Rather, we place NPAS4 in the role of an activity-dependent modulator that acts on top of this baseline connectivity: when NPAS4 is expressed in response to neuronal activity, it shifts the balance of CCK inhibitory input along the somatodendritic axis, increasing somatic and decreasing dendritic CCK synaptic strength [1,2]. The question is therefore not how CCK synapses are established in the absence of NPAS4, but rather how experience-dependent activity uses NPAS4 to fine-tune the distribution of those synapses and it is this fine-tuning that our study links to the precision of in vivo spatial and temporal coding.

      Methodological Clarifications:

      Clarification on how stability analysis was performed. Reviewer #2 requested additional analysis for the stability results: “A control analysis using a fixed reference window around the original peak, rather than re-identifying the peak each epoch, would help distinguish a genuine plasticity-like shift from instability driven by noise.” We wish to clarify that this is precisely the methodology that was used in the manuscript. For the stability analysis shown in Figures 4C-E, the activity was aligned to the peak activity in epoch 1 such that 0 always represents the location of the peak in epoch 1. This approach allows us to identify how that activity differs in subsequent epochs, namely whether it has shifted relative to the activity in epoch 1. We will make this more clear in the results and methods sections.

      Request for Ai32 control. Reviewer #2 made the point that “The comparison throughout the manuscript pits Cre+ ChR2+ neurons (NPAS4 KO) against neighboring non-transduced neurons (WT). This is internally elegant, but leaves open the possibility that part of the phenotype arises from chronic ChR2 expression or constitutive Cre activity rather than from NPAS4 loss, especially given that most of the readouts are subtle.” We agree this would be the ideal control and regret that it is no longer experimentally feasible, as the laboratory in which these experiments were conducted is no longer operating. However, we believe several features of the existing dataset make a ChR2 or Cre artifact unlikely. First, the effects of chronic ChR2 expression are not known to produce the specific pattern of phenotypes we observe in particular the redistribution of somatic versus dendritic inhibition, which is recapitulated independently in acute slice recordings from animals that did not undergo optotagging procedures (Figure 1C). Second, the phenotype we report is internally coherent across multiple independent metrics: place field size, stability, signal-to-noise ratio, theta coupling, and phase precession all shift in the same direction, in a manner consistent with a specific change in inhibitory synaptic balance rather than a nonspecific effect of transgene expression. Third, the sparse nature of the Cre expression means that KO and WT neurons share the same local network, same LFP, and same behavioral context any network-level effect of Cre or ChR2 would be expected to affect both populations similarly. We will add a discussion of these points to the manuscript.

      PSTH clarification (unit of opto-response). To quantify the opto-response, we treated each light-on + light-off period (a total of 2 seconds) as the one trial. We aligned the trials by the light-on period, binned the spikes by 1 msec bins, and then summed the responses across trials to produce a histogram. From this histogram we found the maximum response during light off (e.g. the 1 msec bin with the greatest response which should be reported as number of spikes). We subtracted this from the maximum response during light on. Thus, the unit of opto-response should be spike counts. We will clarify this in the text and figures.

      Use of male mice. Reviewer #1 rightfully pointed out that this study only used male mice. In this study, we only used mice that were larger than 20 grams to ensure the mice could carry the weight of the implanted drives while performing the behavior. As this genetic line of mice is on the smaller size, only male mice were above this weight threshold. Importantly, slice work conducted in the Blood good lab has not identified sex differences in NPAS4 phenotypes [3,9]. Future studies would benefit from the use of both male and female mice. We will state this more explicitly in the text and expand on the potential implications of excluding female mice from our study.

      Future planned changes to manuscript:

      As the reviewers suggested, we intend to add the following analyses and make the following changes to the manuscript:

      Stratify key analyses (stability, theta coupling, phase precession) by FR to determine whether there is a dependency on the firing rate of cells.

      Apply hierarchical bootstrapping and add per-animal color-coding to supplementary figures to assess animal-level variability and protect against pseudoreplication.

      Add a circular-linear phase-position correlation analysis as an additional quantification of phase precession strength, complementing the existing slope-based analysis.

      Improve discussion around the temporal phenotype being downstream of the spatial one.

      Tighten mechanistic framing in the Discussion to more clearly distinguish what is demonstrated in this study from what is inferred from prior work, and to acknowledge the contributions of other inhibitory cell types.

      Minor changes and figure clarifications as noted by reviewers.

      Outside of the scope of this study or unable to be performed:

      There were several recommendations or points that the reviewers brought up that we do not have the resources to address. Nevertheless, we appreciate the reviewers noting these.

      SE control (as discussed above)

      Ai32 control (as discussed above)

      Behavioral consequences of NPAS4 knockout and the effects on learning and memory • Ripple analysis

      Drift observed in E4 and what this might look like over larger timescales

      Comparison between male and female mice to determine whether there are sex-dependence differences

      In conclusion, the reviewers recognized this as a well-designed and internally consistent study. We believe that many of the critiques including the request for a standard housing control, questions regarding the extent of NPAS4 expression across the pyramidal cell population, and points about the timeline of NPAS4 expression and synaptic reorganization are addressed by the clarifications provided in this response. We agree with many of the suggested analytical and textual changes and look forward to incorporating those into the revised manuscript.

      References:

      (1) Heinz, D. A., Cui, W., Cooper, K. L. & Bloodgood, B. L. Experience-induced NPAS4 reduces dendritic inhibition from CCK+ inhibitory neurons and enhances plasticity. J. Neurophysiol. 134, 361–371 (2025).

      (2) Hartzell, A. L. et al. NPAS4 recruits CCK basket cell synapses and enhances cannabinoid-sensitive inhibition in the mouse hippocampus. Elife 7, (2018).

      (3) Bloodgood, B. L., Sharma, N., Browne, H. A., Trepman, A. Z. & Greenberg, M. E. The activity dependent transcription factor NPAS4 regulates domain-specific inhibition. Nature 503, 121–125 (2013).

      (4) Sharif, F., Tayebi, B., Buzsáki, G., Royer, S. & Fernandez-Ruiz, A. Subcircuits of deep and superficial CA1 place cells support efficient spatial coding across heterogeneous environments. Neuron 109, 363–376.e6 (2021).

      (5) Quirk, C. R. et al. Precisely timed theta oscillations are selectively required during the encoding phase of memory. Nat. Neurosci. 24, 1614–1627 (2021).

      (6) Ramamoorthi, K. et al. Npas4 regulates a transcriptional program in CA3 required for contextual memory formation. Science 334, 1669–1675 (2011).

      (7) Chiaruttini, N. et al. ABBA+BraiAn, an integrated suite for whole-brain mapping, reveals brain-wide differences in immediate-early genes induction upon learning. Cell Rep. 44, 115876 (2025).

      (8) Früh, S. et al. Neuronal Dystroglycan Is Necessary for Formation and Maintenance of Functional CCK-Positive Basket Cell Terminals on Pyramidal Cells. J. Neurosci. 36, 10296–10313 (2016).

      (9) Lin, Y. et al. Activity-dependent regulation of inhibitory synapse development by Npas4. Nature 455, 1198–1204 (2008).

    1. eLife Assessment

      This revised study presents valuable findings implicating nuclear export in the regulation of protein condensate behaviour and TDP-43 phase behaviour, suggesting a link to pathogenic aggregation in ALS/FTD. The work contains several observations that will be of interest to the field; however, the underlying mechanistic links proposed by the authors remain insufficiently supported by the current data. The research relies extensively on synthetic, non-physiological protein variants and a homozygous disease model, with limited mechanistic validation, leaving many of the conclusions largely correlative. Thus, despite its technical strengths, the findings presented are currently incomplete, and while the results are invaluable to the field, these do not provide sufficient evidence to substantiate claims about the direct role of nuclear export in pathological protein aggregation and disease.

    2. Reviewer #1 (Public review):

      This revised manuscript represents a partial response to the concerns raised in the first round of review. The authors have made one genuine mechanistic addition in the form of the semi-permeabilized cell reconstitution assay, removed the most overreaching conclusions regarding the contribution of cytoplasmic TDP-43 aggregation to disease, and made several minor presentational improvements. However, the central weaknesses of the original submission remain substantially unaddressed. The exclusive reliance on non-physiological TDP-43 variants, the incompletely resolved mechanism linking XPO1 to TDP-43 phase behavior, and the limited organoid validation continue to limit confidence in the major claims. The authors have, in several instances, responded by removing contested data rather than by providing the additional evidence that was requested.

      (1) The justification for the 2KQ acetylation-mimetic system remains inadequate.

      The authors respond to the concern about the non-physiological nature of the 2KQ mutant by citing published evidence that TDP-43 acetylation occurs in ALS patient spinal cord and is upregulated under oxidative and proteotoxic stress conditions. While these references are real and support the relevance of acetylation as a pathological post-translational modification, they do not resolve the central concern: there is no quantification of how much endogenous TDP-43 is acetylated at the specific lysine residues mimicked by 2KQ in degenerating human neurons, and no evidence that the degree of RNA-binding disruption imposed by the double glutamine substitution is ever achieved by endogenous acetylation in vivo. The 2KQ mutant eliminates RNA binding essentially completely, whereas physiological acetylation events are graded, reversible, and likely partial. The response conflates the existence of TDP-43 acetylation as a phenomenon with validation that 2KQ is a physiologically accurate model of that phenomenon. None of the new experiments address the request to test whether wild-type TDP-43 expressed at near-physiological levels, or a bona fide heterozygous ALS-linked TARDBP mutant in iPSC-derived neurons, responds to XPO1 modulation in a qualitatively similar fashion. Until this is shown, the mechanistic conclusions of this paper remain constrained to a highly artificial overexpression system and cannot be extrapolated to physiological or pathological TDP-43 biology with confidence.

      (2) The homozygous K181E organoid model is still not adequately justified, and no heterozygous comparison has been provided.

      The authors acknowledge that the homozygous background is "more sensitive for detecting phospho-TDP-43" and argue that homozygous conditions are commonly used in experimental TDP-43 research. However, the critical issue is not whether homozygous models are used in general, but whether the homozygous background specifically alters the relative contribution of cytoplasmic aggregation versus nuclear RNA-processing dysfunction in this study. In a homozygous K181E model, both alleles produce an RNA-binding-defective TDP-43, meaning that every molecule of endogenous TDP-43 in the cell is dysfunctional. This is categorically different from the patient situation in which one wild-type allele is present, and it may substantially exaggerate nuclear loss-of-function relative to cytoplasmic gain-of-function phenotypes. The authors have not performed the requested comparison with heterozygous K181E/+ organoids, nor have they acknowledged that the organoid genotype itself could bias the interpretation of what KPT-276 treatment rescues. Given that the organoid section is now the sole in-disease-model validation of the XPO1 mechanism, this limitation is more consequential than it was in the original submission.

      (3) The new semi-permeabilized cell data is a genuine contribution, but the mechanistic interpretation remains insufficiently constrained.

      The development of the streptolysin O semi-permeabilized cell reconstitution system is the most substantive new addition to this revision. The finding that LMB-stabilized anisosomes resist cytosol washout but dissolve upon RNase T1 treatment is interesting and provides a plausible indirect mechanism: XPO1 inhibition retains nuclear RNA, and this elevated nuclear RNA availability contributes to maintaining the liquid LLPS state of the TDP-43 2KQ condensate. This is a meaningful mechanistic advance and deserves credit. However, several important limitations of this new data are not adequately discussed. First, RNase T1 degrades single-stranded RNA globally during permeabilization, so the experiment does not identify which specific RNA species stabilize the anisosome, nor whether these are pre-mRNA splicing intermediates, mature mRNA, non-coding RNA, or another class. Second, the same nuclear export blockade that retains RNA will also retain the nuclear concentrations of many RNA-binding proteins, splicing factors, and other XPO1-dependent cargos. The RNase T1 experiment does not exclude the possibility that the relevant effect is mediated by an RNA-binding protein whose nuclear concentration increases upon LMB treatment and which, upon RNase digestion, can no longer engage TDP-43 or the anisosome shell. Third, the permeabilized cell system is by definition not intact and has lost cytosolic factors; whether the RNA-dependent stabilization of anisosomes operates in the same way in intact cells during physiological or pathological nuclear export perturbation is an assumption, not a demonstrated fact. The authors should more carefully frame these data as hypothesis-generating and explicitly note these alternative interpretations in the Discussion.

      (4) The conceptual asymmetry between XPO1 inhibition and XPO1 overexpression phenotypes is not resolved by the new mechanism.<br /> The paper continues to present two XPO1 perturbation phenotypes that are difficult to reconcile within a single mechanistic model. XPO1 inhibition enlarges anisosomes, maintains their liquid character by FRAP, and retains them in the nucleus. XPO1 overexpression also enlarges TDP-43 puncta, but these are FRAP-impaired, gel-like, and appear in the cytoplasm. The RNA-retention model proposed by the new semi-permeabilized data explains why XPO1 inhibition stabilizes the liquid state, but it does not explain why XPO1 overexpression drives the opposite outcome: gel-like hardening and cytoplasmic redistribution. If increased nuclear RNA availability is the key variable downstream of XPO1 inhibition, then XPO1 overexpression would be expected to decrease nuclear RNA and thereby destabilize anisosomes toward dissolution or hardening. The paper does not test whether nuclear RNA levels are indeed altered by XPO1 overexpression, nor whether the cytoplasmic gel-like puncta seen in XPO1-overexpressing cells are RNA-poor relative to control anisosomes. The revised Discussion does not engage with this asymmetry in a satisfying way, and the figure model remains qualitative. A quantitative or at least semi-quantitative model that accounts for both arms of the XPO1 perturbation is needed.

      (5) The removal of RNA-seq data weakens rather than strengthens the organoid section.

      The authors have removed the bulk RNA-seq analysis from the revised manuscript in response to concerns that the modest transcriptional rescue was being over-interpreted. While the decision to remove over-interpretation is appropriate, the result is that the organoid section now rests entirely on pTDP-43 immunostaining as its sole readout. The revised paper thus uses reduction in immunofluorescent pTDP-43 puncta in homozygous K181E organoids as the only evidence that nuclear export inhibition mitigates TDP-43 proteinopathy in a disease-relevant context. This is a weaker evidentiary base than before the revision, not an improvement. The originally requested more sensitive orthogonal readouts, including biochemical fractionation for SDS-insoluble TDP-43, filter-trap assays, or RNA aptamer-based detection of TDP-43 aggregates, remain absent. Without at least one additional independent measure confirming that cytoplasmic TDP-43 aggregation is genuinely reduced rather than simply rendered antigenically undetectable, the organoid conclusion is not adequately supported. At minimum, the authors should provide total and cytoplasmic TDP-43 fractionation data from organoid lysates to corroborate the immunostaining result.

      (6) No functional neuronal readout has been provided for the organoid model.

      The organoid section now makes the claim that "nuclear export is required for the formation of p-TDP-43-containing aggregates in a disease-relevant organoid model," but no measure of neuronal health, integrity, or function is reported in association with this. Even a simple assessment of neuron survival by TUJ1 or MAP2 quantification, neurite complexity, or cleaved caspase-3 staining before and after KPT-276 treatment would substantially strengthen the biological significance of the pTDP-43 reduction. The current data establish a pharmacological effect on a pathological marker but do not demonstrate that this has any consequence for neuronal biology in the organoid, which is what the disease-relevance framing implies.

      (7) The abstract and title continue to overstate the mechanistic conclusions.<br /> Despite the stated intent to reframe the study as a screening study and to temper the conclusions, the revised abstract retains the language: "These findings establish nuclear export as a key regulator of TDP-43 phase transitions and define a mechanistic framework that links altered nuclear transport and phase dynamics to TDP-43 aggregation potential." Similarly, the Discussion still states: "a particularly compelling aspect of our study is the discovery that the nuclear export receptor XPO1 governs TDP-43 liquid-to-solid transitions and subcellular localization." The word "governs" and the phrase "establish nuclear export as a key regulator" are not warranted by data that derive entirely from an overexpressed acetylation-mimetic mutant in a colon cancer cell line and a homozygous K181E organoid model. A more accurate framing would describe these findings as identifying nuclear export as one of several cellular processes that modulate TDP-43 phase behavior in a sensitized model system, with an indirect RNA-mediated mechanism that remains to be defined at the molecular level. The title change from "governs" to "modulates" is appreciated but does not extend into the abstract and Discussion, where the strong causal language persists.

      (8) Individual siRNA knockdown validation for XPO1 has not been provided.

      The authors argue that validation with 6 independent siRNAs across two rounds of screening, combined with convergent pharmacological data, is sufficient to establish XPO1 as a genuine hit. While the convergence of chemical and genetic evidence is reassuring, the specific request was for protein-level confirmation of XPO1 knockdown efficiency in the DLD1 TDP-43 2KQ cells used for mechanistic follow-up, together with demonstration that the anisosome phenotype is specifically caused by loss of XPO1 and not by off-target effects. This is a straightforward experiment, and its absence is particularly notable given that the entire mechanistic XPO1 narrative hinges on this specificity. At minimum, an immunoblot confirming XPO1 protein depletion in cells treated with the siRNA pool identified in the screen, in the same cell background and induction conditions as the follow-up experiments, should be provided.

      (9) The identity of XPO1-dependent cargos that regulate anisosome dynamics remains entirely unknown.

      The authors acknowledge that XPO1 does not directly bind TDP-43 and that the mechanism is likely indirect. The new RNA data provides one plausible indirect pathway. However, the possibility that one or more specific RNA-binding proteins or splicing factors, whose nuclear levels rise upon XPO1 inhibition, are the proximate drivers of anisosome stabilization has not been addressed. This matters because if the relevant mechanism operates through a specific cargo rather than bulk RNA retention, the model for how nuclear export connects to TDP-43 aggregation in disease would be fundamentally different. The authors decline to pursue adaptor identification on grounds of scope, which is a defensible position for future work. However, the framing should explicitly state that the current data cannot distinguish between bulk RNA retention and cargo-specific effects, and that the conclusion that nuclear export modulates TDP-43 phase behavior via RNA accumulation is a working hypothesis supported by but not proven by the RNase T1 experiment.

      Minor remaining issues.

      The number of independent iPSC clones and organoid batches used for the KPT-276 treatment experiment is now stated as two batches per condition, which is minimal for a 3D organoid study and does not fully address the concern about clone-level variability. Ideally, organoids from at least two independently derived isogenic clones per genotype would be used. The mCherry overexpression control added in Supplemental Figure 4 is a useful addition and is acknowledged. The immunoblotting confirmation that drug treatments do not alter total TDP-43 levels addresses a prior concern adequately. The addition of the sentence noting that anisosomes have not been validated in human patient samples is appreciated and appropriate. Statistical detail has been improved in figure legends. These minor improvements are noted positively but do not compensate for the major unresolved concerns above.

    3. Reviewer #2 (Public review):

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. They further strengthen their findings by introducing an additional model system, a semi-permeabilized in vitro assay, which provides mechanistic evidence that XPO1 activity prevents anisosome dissolution by retaining nuclear RNAs. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. Some minor concerns remain, mostly about experimental design of the newly added data.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      (6) The authors extend their study by incorporating a semi-permeabilized in vitro system, which provides compelling evidence that inhibition of nuclear export promotes the retention of nuclear anisosomes, an effect driven by the accumulation of nuclear RNAs.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      Weaknesses:

      (1) The manuscript has significantly improved with the revisions. Some experimental procedures and method details, as well has statements remain incompletely described:

      a) What is the smear in Figure S1 after VLX treatment?

      b) The authors state that "The reduction in TDP-43 signal was not due to protein elimination.", however no data is provided to support that statement.

      c) The authors state that "TDP-43 shifts from phase-separated state to a soluble state ...", however no data is provided to support that statement.

      d) Why did the authors choose cow lover cytosol for this study?

      e) The experimental setup for supplementing with cytosol/ATP/GTP is unclear. A more detailed schematic would be helpful to understand at what stage in the experiment these factors were added. Which step of the protocol was performed at 37 {degree sign}C, which is indicated in the figure schematic but not described in the methods.

      f) In the organoid model, the authors mention that they observe similar levels of total TDP-43, however they do not provide quantification. Instead, they provide a graph that shows highly significant changes in nuclear TDP-43, which was not addressed in the text.

      Additionally, some questions remain unclear:

      (1) The anisosomes induced by ATP/GTP or cytosol are insufficiently characterized. It remains unclear whether these structures correspond to canonical ring-shaped anisosomes, and whether they exhibit dynamic (liquid-like) or more static (gel-like) properties.

      (2) The contribution of the cytosol and ATP/GTP supplementation experiments to the overall narrative is unclear. While the findings are intriguing, their interpretation within the context of the study is not well articulated. In particular, the rationale for including cytosol is not sufficiently justified, given that ATP/GTP alone induces a pronounced effect, whereas cytosol alone does not.

      (3) The authors should address why endogenous XPO1 does not co-localize with anisosomes, whereas overexpressed XPO1 does. This raises the possibility that the observed co-localization may be an artifact of non-physiological protein levels, which should be discussed.

      (4) The iPSC-based model remains insufficiently characterized. While the authors propose that this system recapitulates the accumulation of liquid and solid aggregates resembling anisosomes, it is unclear whether this phenotype is robustly observed and whether KPT treatment effectively modulates it.

      (5) The rationale for the selected treatment durations is unclear, and the timing appears inconsistent across experiments (ranging from 3 to 16 hours), including within experiments involving the same compound. This variability should be justified or standardized.

      (6) Several figure legends require clarification:

      a) In the section stating "Collectively, our results suggest that the stability and dynamics of anisosomes are modulated by XPO1-mediated nuclear export ...", the cited figure appears to be incorrect. This should refer to Figure 5L rather than Figure 5J.

      b) Figure 1B: Please specify the number of replicates per concentration, the number of cells analyzed, and the model used for regression analysis. Additionally, the legend indicates a treatment duration of 15 hours, whereas Figure 1A states 24 hours.

      c) Figure 2G: The authors state "7 anisosomes per condition," but the graph displays only 4-6 data points. Please clarify what each data point represents.

      d) Figures 3B and 3G: Please clarify whether a defined threshold was used to determine a "reduction in anisosome number."

      e) Figure 4B: These do not represent biological replicates, as all samples derive from a single cell line; rather, they constitute independent experimental replicates.

      f) Figures 5B and 5H: The legend states "n = 3 biological repeats," but the number of data points shown appears higher. Please clarify.<br /> g) Figures 5K, 6C, and 6E: "Mean Fluorescence Intensity (MPI)" should be corrected to "MFI."

      h) Figure 6C: Please include the number of cells analyzed and provide relevant statistical measures (e.g., R<sup>2</sup>, p-value).

      i) Figure 6D: The experimental timeline is unclear. Please specify the duration of incubation and the timing of each step.

      j) Figure 7B: Improved labeling is needed (e.g., clarification of "mean spot volume") to better align with the figure legend.

    4. Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      Comments on revised version.

      The manuscript has been improved with more data and clarification. The RNase T1 treatment experiment suggests that RNA is required for anisosome integrity. However, this does not directly demonstrate LMB increases nuclear RNA availability as changes in protein composition or other RNA-dependent mechanisms may also contribute. The conclusion and discussion need to be edited to consider these alternative scenarios. Overall, as most of the evidence remains indirect, the manuscript should avoid overinterpretation regarding the mechanisms underlying TDP-43 phase transition and aggregation.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, the authors use a doxycycline-inducible DLD1 cell line expressing a Clover-tagged RNA-binding-defective TDP-43 2KQ mutant that forms nuclear "anisosomes" (TDP-43 shell with HSP70 core) to carry out a small-molecule screen using the LOPAC 1280 library to identify compounds that reduce anisosome number or shift their morphology and dynamics. They also conducted a genome-wide siRNA screen to identify genetic modifiers of anisosome formation and dynamics. From these screens, the authors identify pathways in RNA splicing, translation, proteostasis (proteasome and HSP90), and nuclear transport, including XPO1. They then focus on XPO1 as their primary hit. Pharmacological inhibition of XPO1 using KPT-276, Verdinexor, and Leptomycin B reduces anisosome number while enlarging remaining condensates, which retain liquid-like behavior by FRAP and fusion assays. XPO1 overexpression causes fewer, enlarged TDP-43 puncta, including cytoplasmic puncta, with little or no FRAP recovery, interpreted as gel or solid-like aggregates. Anisosome induction reduces detectable nucleoplasmic XPO1 staining. Finally, the authors examine a homozygous TDP-43 K181E iPSC-derived forebrain organoid model, showing increased cytosolic pTDP-43 in K181E/K181E organoids compared to wild-type controls. Chronic low-dose KPT-276 reduces cytoplasmic pTDP-43 without changing total TDP-43 levels. Bulk RNA-seq shows only a modest fraction of dysregulated genes in K181E/K181E organoids are rescued by KPT-276. They conclude that nuclear export, via XPO1, is a key regulator of TDP-43 liquid-to-solid phase transitions and that cytoplasmic aggregation per se may contribute only modestly to TDP-43 proteinopathy, with RNA-processing defects being dominant.

      We thank the reviewer for carefully summarizing our study.

      The study presents well-executed chemical and genome-wide siRNA screens in a DLD1 TDP-43 2KQ anisosome model and follows up on nuclear transport, particularly XPO1, as a modulator of TDP-43 phase behavior and cytoplasmic aggregation. The screens are impressive in scale, and the microscopy and fluorescence recovery after photobleaching (FRAP) work is technically strong. However, the central mechanistic and disease-relevance claims are not yet sufficiently supported. There are major concerns about the heavy reliance on non-physiological, RNA-binding-defective, and acetylation-mimetic TDP-43 (2KQ) and a homozygous TDP-43 K181E organoid model. An underdeveloped and partly contradictory mechanistic link exists between XPO1 and TDP-43 phase transitions in the context of prior work showing TDP-43 is not a canonical XPO1 cargo. The paper also appears to overinterpret organoid data to conclude that cytoplasmic TDP-43 aggregation plays only a minor role in pathology, based largely on pTDP-43 antibody staining with limited sensitivity and relatively modest rescue readouts. A deeper mechanistic analysis and additional, more physiological validation are needed for this to reach the level of rigor and impact implied by the title and abstract. The work feels screen-rich but conceptually underdeveloped, with key claims outpacing the data. A major revision with substantial new data and tempering of conclusions is warranted. I outline several problematic areas below:

      (1) The central mechanistic discoveries are derived almost entirely from a DLD1 colon cancer cell line overexpressing an RNA-binding-defective, acetylation-mimetic TDP-43 2KQ mutant and homozygous TDP-43 K181E iPSC-derived organoids. Both systems are far from physiological. The 2KQ mutation is a synthetic double lysine-to-glutamine mutant originally designed to mimic acetylation and disrupt RNA binding. In this study, essentially all cell-based mechanistic data on phase behavior, screens, and XPO1 effects rely on 2KQ. Yet there is no quantification of how much endogenous TDP-43 is acetylated in degenerating human neurons, nor whether a 2KQ-like acetylation state is ever achieved in vivo. It is not established that the phase behavior of 2KQ recapitulates the physiological or pathological phase behavior of wild-type TDP-43 or genuine disease-linked mutants, which may retain partial RNA binding and different post-translational modification patterns. As a result, it is difficult to know whether the modifiers identified here regulate a highly artificial 2KQ condensate or physiologically relevant TDP-43 condensates. To address this concern, the paper would benefit from quantifying endogenous TDP-43 acetylation at the relevant lysines in control and ALS/FTD patient tissue or more disease-proximal models such as heterozygous TARDBP mutant iPSC neurons, which would justify the focus on an acetyl-mimetic mutant. Key phenomena, including XPO1 dependence of phase behavior, effects of proteasome and HSP90 inhibition, and effects of splicing and translation inhibitors, should be tested for wild-type TDP-43 expressed at near-physiological levels and for one or more bona fide ALS/FTD-linked TARDBP mutants that are not acetyl mimetics. At a minimum, the authors should show that endogenous TDP-43 in neuronally differentiated cells exhibits qualitatively similar responses to XPO1 modulation, rather than exclusively relying on DLD1 2KQ overexpression.

      Acetylation of endogenous TDP-43 was reported by several studies. Although it occurs at low levels under normal conditions, TDP-43 acetylation is upregulated under stress conditions (e.g. oxidative stress and proteotoxic stress) (PMID: 25556531; PMID: 28724966). Importantly, Cohen et al. reported the identification of acetylated TDP-43 in ALS patient spinal cord (PMID: 25556531), while Yu et al. showed that endogenous wildtype TDP-43 undergoes demixing when neurons were treated with either a deacetylase inhibitor or proteasome inhibitor (PMID: 33335017). These studies also show that acetylated TDP-43 is defective in RNA binding and more prone to aggregation. Furthermore, ectopic expression of acetylated TDP-43 mimetics in cells and mice induces cellular defects similar to those observed in disease models (PMID: 28724966). Thus, our findings, based on previously established TDP-43 mimetics, should provide valuable information regarding the phase regulation of a disease-relevant TDP-43 mutant. We have included more background information to justify the use of TDP-43 acetylation mimetics in the introduction.

      (2) The organoid model is based on a homozygous K181E knock-in line. However, in patients, TARDBP mutations are overwhelmingly heterozygous. Homozygosity is thus a severe, arguably non-physiological sensitized background that may exaggerate nuclear RNA mis-splicing and phase defects and alter the relative contribution of cytoplasmic aggregation versus nuclear loss-of-function. In addition, it is not fully clear from this manuscript whether the structures in K181E organoids are bona fide anisosomes as defined in Yu et al. 2021, characterized by HSP70-enriched central liquid cores with TDP-43 shells and similar FRAP and fusion behavior to anisosomes in the DLD1 model. At present, the organoid section is framed as validation of "anisosome-bearing organoids," but the figures in this manuscript mainly show pTDP-43 puncta and total TDP-43 immunostaining, without detailed structural or biophysical characterization. The authors should explicitly compare heterozygous K181E/+ organoids or another heterozygous TARDBP mutant line with homozygous K181E/K181E organoids to assess whether XPO1 inhibition has similar effects in a genotype that more closely resembles patient genetics. They should provide direct evidence that the K181E condensates in organoids are anisosomes through HSP70 core immunostaining, three-dimensional reconstruction, and FRAP measurements, and clarify whether KPT-276 is acting on anisosome-like structures or more generic cytoplasmic aggregates or puncta. Without this, the leap from a DLD1 2KQ cancer cell model to human ALS/FTD-relevant neurons is not convincingly supported.

      The reviewer is correct that the use of homozygous K181E organoids generates a background that is more sensitive for detecting phospho-TDP-43. The goal was to test whether XPO1 inhibition mitigates the phosphorylation of a TDP-43 disease mutant. For this purpose, we believe that our experimental setup is suitable. We agree that we should not extrapolate the result to over emphasize on its disease connection. We have revised the paper to tone down this section. We also remove the RNAseq data as it is not essential for our conclusions.

      It is also noteworthy that TDP-43 disease mutations are usually loss-of-function alleles. Although heterozygous background is sufficient to induce disease phenotype in aged humans, heterozygous background in experimental settings is usually unable to generate severe defects. Thus, it is quite common to study TDP-43 disease-related defects in homozygous knockout or RNAi-mediated depletion conditions (e.g. PMID: 35197626; 41120751; 38277467).

      Regarding the immunostaining signals in K181E organoids, we did not report them as anisosomes. As documented in the literature, p-TPD-43 is widely used as a marker to indicate pathological TDP-43 aggregation. P-TDP-43 is enriched in pathological aggregates in human ALS and FTD patients, colocalized with other aggregation signatures such as ubiquitin and other aggregation-prone proteins in the cytoplasm (PMID: 36008843), and is being used as a diagnostic marker for neurodegeneration (PMID: 31661037). The characterization of K181E organoid is reported in a pre-print by Zhang Q. et al., 2026 (PMID: 41292965), which is currently under revision for Science Advances. In Fig. 1I of this manuscript, we confirmed the cytosolic localization of p-TDP-43 in cells that were isolated from K181E organoids. In the current manuscript, Figure 7 is to show that nuclear export inhibition mitigates the accumulation of p-TDP-43 in a brain-like tissues. We revise the subheading and the corresponding text to avoid the confusion.

      (3) The title and framing assert that "nuclear export governs TDP-43 phase transitions." However, prior studies such as Pinarbasi et al. 2018 and Duan et al. 2022 indicate that TDP-43 is not a canonical XPO1 cargo and that its export is largely passive, with active nuclear import being the dominant determinant of nuclear localization. The authors cite these studies but still position XPO1 as a central, quasi-direct regulator. The data presented are largely correlative or based on pharmacologic manipulation and overexpression in an overexpression mutant background, with no direct evidence that XPO1 engages TDP-43 in a specific, regulated manner. Even if XPO1 does not engage WT TDP-43, it could still engage the 2KQ variant, which needs to be tested.

      We did not mean to conclude or imply that the regulation of TDP-43 by XPO1 is direct. In fact, we explicatively mentioned on page 8 of the original manuscript that the regulation is likely indirect and mediated by other factors. The sentence reads as “Since XPO1 does not bind TDP-43 directly (Pinarbasi et al., 2018), additional factors might link XPO1-mediated nuclear export to TDP-43 nuclear egression.”

      We now add new data in Figure 6, showing that in an in vitro reconstitution assay using semi-permeabilized cells, LMB treatment significantly stabilizes anisosomes in an RNA dependent manner. This new data suggests that XPO1 inhibition leads to increased nuclear RNA availability, which indirectly favors anisosome assembly and maturation (see discussion). We believe that this new finding has provided significant new insight into how nuclear transport modulates TDP-43 phase behavior. We have revised the title, the abstract and changed the framing according to the reviewer’s suggestion.

      (4) The XPO1 perturbations yield somewhat confusing phenotypes. XPO1 inhibition using Leptomycin B, KPT-276, and Verdinexor reduces anisosome number and enlarges remaining anisosomes, which remain liquid-like by FRAP recovery and fusion assays and stay nuclear. XPO1 overexpression causes fewer, enlarged puncta, but these are FRAP-impaired (gel-like) and redistribute to the cytoplasm. Thus, both decreased and increased XPO1 activity reduce anisosome number and enlarge puncta, but with opposite phase behaviors and subcellular localizations. The model presented in Figure 5L is relatively qualitative and does not resolve these issues. Moreover, XPO1 inhibition globally impairs nuclear export of many cargos and profoundly alters the nuclear environment, transcription, RNA processing, and chromatin. It is therefore difficult to conclude that the observed effects are specific to TDP-43 phase regulation as opposed to secondary consequences of broad nuclear export blockade.

      The reviewer correctly summarizes our data and interpretation: XPO1 loss-of-function and gain-of-function generate opposite phenotypes regarding TDP-43 phase regulation.

      Regarding the mechanism underlying XPO1-dependent TDP-43 phase regulation, as mentioned above, we developed a semi-permeabilized cell-based assay in which we used the pore-forming toxin streptolysin O to damage the plasma membrane after anisosome induction. We noticed that upon cell permeabilization and cytosol loss, anisosomes were mostly lost (Figure 6B, C). This is probably due to a reversible partition of TDP-43 into a less fluorescent soluble fraction. Supporting this idea, when permeabilized cells were incubated with cytosol plus an energy regenerating system, small puncta containing TDP-43 2KQ could be reformed in an energy dependent manner (Figure 6D, E). Interestingly, in LMB-treated cells, anisosomes remained stable despite cell permeabilization(Figure 3F). Since LMB treatment did not increase TDP-43 nuclear concentration (Supplemental Figure 1), this data suggest that nuclear export inhibition likely alter the nuclear environment to stabilize anisosomes. Indeed, when cells were permeabilized in the presence of a small RNAase, LMB-stabilized anisosomes also collapsed (Figure 6G).

      We now add more discussions on the potential effect of RNA on TDP-43 phase behavior in XPO-1 inhibited cells considering these new findings.

      (5) The authors show that anisosome induction depletes nucleoplasmic XPO1 signal and that mCherry-XPO1 can be seen in some TDP-43 puncta. However, antibody penetration into anisosomes is limited, so XPO1 depletion from nucleoplasm could reflect sequestration in the anisosome shell or core, but this is not demonstrated. There is no demonstration of physical interaction, even indirect interaction, between XPO1 and TDP-43 or a defined adaptor, nor identification of a specific mutant of XPO1 that selectively disrupts this putative interaction while preserving other functions. The known TDP-43 NES has been shown to be weak and not a functional XPO1-dependent NES in multiple studies. If XPO1 is acting through an adaptor that recognizes 2KQ or K181E specifically, that by itself would bring into question the generality of the mechanism for wild-type TDP-43.

      We agree that our data does not demonstrate an interaction between XPO1 and TDP-43. Considering our new data (mentioned above), it is possible that the effect of anisosome induction on endogenous XPO1 localization is also mediated by RNA. We now mention more explicitly that the regulation of TDP-43 by XPO1 is likely indirect (Page 8). We have revised our paper to separate any speculative statements from the data, and also discussed the possibility of alternative interpretations.

      (6) To support a mechanistic claim that nuclear export governs TDP-43 phase transitions, more targeted evidence is needed. The authors should test whether siRNA knockdown or CRISPR interference of XPO1 in the DLD1 2KQ model reproduces the effects seen with Leptomycin B and KPT-276, including FRAP and fusion phenotypes, and verify on-target effects by rescue with an siRNA-resistant XPO1 construct. They should demonstrate that canonical XPO1 cargos behave as expected under the inhibitor conditions used, as a positive control, and that the concentrations used are not grossly toxic. They should attempt to identify or at least constrain candidate adaptors that might enable XPO1-dependent export of TDP-43 through proteomic analysis of XPO1 co-purifying with 2KQ condensates or loss-of-function studies of candidate adaptors from the siRNA screen. Finally, they should test whether a TDP-43 mutant that cannot bind the proposed adaptor still responds to XPO1 manipulation.

      The anisosome enlargement phenotype upon XPO1 depletion was seen in our siRNA screens, which was identified by machine-based image analyses using 6 different siRNAs. This, together with the chemical inhibition experiments, demonstrate that the phenotype is specifically caused by XPO1 inactivation.

      When characterizing the effect of XPO1 inhibition on anisosome dynamics, we preferred chemical inhibitor because the effect is acute, and therefore less likely to be secondary.

      Regarding the inhibitor concentration, according to the literature, Leptomycin B was commonly used at 50-200 nM. We chose 200 nM to ensure a quick and complete inhibition of XPO1-mediated nuclear export (see Figure 3 in PMID: 9628873). This dose is also well tolerated by our cells.

      We did not suggest any specific adaptor that mediates XPO1 interaction with TDP-43. Whether there is an adaptor, and if so, the identity of such adaptor is out of the scope of this study. We revise our paper on page 8-9 to clarify these points.

      (7) Even with these data, what is currently shown is that global modulation of nuclear export capacity can alter the phase behavior and localization of a highly overexpressed RNA-binding-defective TDP-43 mutant and of K181E in organoids. This is important, but it is weaker than asserting that XPO1 directly governs TDP-43 phase transitions in physiological contexts. The title, abstract, and Discussion should be tempered to reflect that nuclear export is one of several pathways, alongside RNA splicing, translation, and proteostasis, that influence TDP-43 phase states in this model, and that the specific mechanism and cargo relationship between XPO1 and TDP-43 remain unresolved and may be indirect.

      We have revised the title, abstract, and main text to temper our conclusions.

      (8) The authors conclude that cytoplasmic TDP-43 aggregation plays only a modest role in TDP-43 proteinopathies because in homozygous K181E organoids, chronic KPT-276 treatment almost abolishes cytoplasmic pTDP-43 puncta, yet bulk RNA-seq shows only a relatively small fraction of dysregulated genes are rescued. There are several issues with this inference. Relying primarily on pTDP-43 antibody staining to define cytoplasmic TDP-43 aggregation is limiting. pTDP-43 antibodies label only phosphorylated species and may miss non-phosphorylated, oligomeric, or amorphous TDP-43 species that could still be toxic. Different pTDP-43 antibodies vary in epitope accessibility depending on aggregate conformation and subcellular location. More sensitive approaches, such as high-affinity TDP-43 RNA aptamer probes developed by Gregory and colleagues, biochemical fractionation for SDS-insoluble and urea-soluble TDP-43, and filter-trap assays, would provide a more quantitative assessment of cytoplasmic aggregation and its reduction by KPT-276. Without these, it is not safe to assume that cytoplasmic aggregation has been eliminated, as opposed to one antigenic subclass.

      We agree with the reviewer that p-TDP-43 may not represent all aggregate species. However, p-TDP-43 antibodies detect the pathologically validated species tightly associated with TDP-43 proteinopatheis. In human ALS and FTD-TDP tissues, cytoplasmic inclusions are strongly immunoreactive for phosphorylated TDP-43 (typically S409/410, as detected here). Additionally, p-TDP-43 immunohistochemistry is a routine diagnostic criterion in neuropathology. For these reasons, we believe that the observation that inhibition of XPO1 significantly reduces p-TDP-43 is a significant finding, as it suggests that inhibition of nuclear transport may rescue TDP-43 proteinopathy. We revised the text on page 9 to better explain the significance of p-TDP-43 staining.

      (9) The treatment window, spanning from day 87 to 122 with 20 nanomolar KPT-276, may be too late or too mild to reverse entrenched nuclear RNA-processing defects, even if cytoplasmic inclusions are cleared. Once widespread cryptic exon inclusion and alternative polyadenylation misregulation are established, many downstream changes may become self-sustaining or only partially reversible. Moreover, XPO1 inhibition will massively rewire nucleocytoplasmic transport of many transcription factors, splicing factors, and RNA-binding proteins. Thus, the lack of full transcriptomic rescue cannot be cleanly interpreted as evidence that cytoplasmic aggregates are only modest contributors. It may instead reflect that nuclear dysfunction is primary and XPO1 inhibition does not correct, and may even exacerbate, certain nuclear defects.

      We agree with the reviewer that the lack of rescue may be caused by some technical issues. We have removed the RNAseq data and the related texts since it is not essential.

      (10) To support a causal statement about the modest contribution of cytoplasmic aggregates, one would want more direct measures of neuronal health and function, such as cell death, neurite complexity, synaptic markers, and electrophysiology before and after KPT-276, not only transcriptomics. A way to selectively reduce cytoplasmic aggregation without globally inhibiting nuclear export would allow comparison of outcomes.

      We have removed the discussion regarding the role of cytoplasmic aggregates in disease.

      (11) Given these caveats, the concluding statements that cytoplasmic TDP-43 aggregation is only a modest contributor should be substantially softened. A more defensible interpretation is that in this homozygous K181E organoid model, chronic global XPO1 inhibition reduces pTDP-43-positive cytoplasmic puncta but only partially normalizes the steady-state transcriptome, suggesting that persistent nuclear RNA-processing defects and other pathways continue to drive pathology.

      We agree with the review and have removed the RNAseq part.

      (12) The screens are a major strength but need more rigorous validation for key hits, especially nuclear transport factors. For the siRNA screen, hits are filtered by anisosome number per nucleus, but there is no direct demonstration in the main text that XPO1 or CSE1L knockdown is efficient at the messenger RNA or protein level. For the highlighted genes, Western blot or quantitative polymerase chain reaction validation and phenotypic rescue would strengthen confidence. For small-molecule hits, it is not systematically shown that anisosome modulation is independent of changes in total TDP-43 2KQ expression or gross toxicity. Translation inhibitors are tested for this, but for many other hits, including proteasome, HSP90, and kinase inhibitors, expression and general nuclear structure should be monitored. Given the reliance on anisosome count as a readout, secondary screens that specifically distinguish changes in TDP-43 expression levels, changes in nuclear morphology or cell cycle, and specific changes in anisosome phase behavior, including FRAP and fusion for top hits, would greatly increase interpretability.

      For the siRNA screen, each positive hit was confirmed by two rounds of screen with 6 independent siRNAs in total. Although we did not validate the knockdown efficiency due to the large number of hits, we routinely include a positive siRNA control in our study (Cell death siRNA), which targets several essential gene. Transfection efficiency was controlled by measuring cell viability after knocking down of these genes. In addition, the identification of XPO1 as a positive regulator of TDP-43 phase behavior was independently validated by our chemical genetic screens with three XPO-1 inhibitors. We feel confident that XPO1 is a key modulator of TDP-43 phase behavior.

      For chemical treatment experiments, the anisosome fusion phenotypes could be detected as early as 5 h post treatment. Given the relatively short treatment, we do not expect a significant change in protein level or toxicity. To alleviate this reviewer’s concern, we performed an immunoblotting experiment to measure the total TDP-43 protein levels in drug-treated cells. Except for VLX, we did not detect any significant changes in the level of TDP-43 after drug treatment (Supplemental Figure 1).

      (13) The classification of condensates as liquid versus gel-like or solid is based almost entirely on FRAP recovery or lack thereof. While FRAP is appropriate, interpretations could be made more robust by including half-region-of-interest bleach controls and assessing mobile fractions and recovery kinetics more quantitatively across conditions. Complementing FRAP with other phase-behavior assays such as sensitivity to 1,6-hexanediol, shape relaxation after deformation, and coarsening behavior over longer timescales would strengthen the analysis. At present, some assignments, such as that XPO1 overexpression drives a gel-like transition, are reasonable but somewhat qualitative.

      In this study, we used two types of FRAP assays. We either bleached TDP-43 within anisosomes or bleached the surrounding TDP-43 molecules(Figure 2). The two complementary methods yield consistent results that allow unambiguously distinguish between TDP-43 LLPS state and gel-like condensation.

      In XPO1-related experiments, the two types of condensates formed by TDP-43 2KQ can be distinguished by several features including their subcellular localization, shape, and the fluorescence recovery kinetics. We feel that these combined data clearly segregate these puncta into two distinct types of assemblies. The proposed half-region-of-interest bleach is technically challenging for small anisosomes under normal conditions. However, whenever possible, (e.g. anisosomes enlarged by Leptomycin B), we did perform both whole anisosome bleach and partial bleach (Figure 5D, I). Both assays demonstrate that TDP-43 in these enlarged anisosomes is highly mobile.

      (14) For the Leptomycin B and KPT-276 experiments in cells and organoids, it would be important to confirm that canonical XPO1 cargo proteins accumulate in the nucleus and that the concentrations used are within a range that is not overtly toxic over the experimental timeframe. Assessing nuclear morphology, chromatin condensation, and general transcriptional activity through global RNA synthesis or key reporter genes would ensure that observed effects are not secondary to severe global nuclear export collapse.

      In Leptomycin B treatment experiments, we carefully chose a dose that was previously validated (see Figure 3 in PMID: 9628873). Based on our DAPI staining, the nuclear morphology appears normal with no abnormal chromosome condensation (Figure 5A). Additionally, in cell line-based experiments, the effect of Leptomycin B on anisosomes was detected 6-8 hours post treatment. The change in global protein synthesis because of RNA changes should be relatively minor at this stage. Indeed, our new immunoblotting experiment showed that LMB treatment did not affect TDP-43 protein level (Supplemental Figure 1). Most importantly, the in vitro semi-permeabilized assay demonstrates a direct role for RNA in stabilizing anisosomes.

      (15) In the organoid section, it is not clear how many independent iPSC clones and organoid batches were used per condition, nor whether batch effects were assessed in the bulk RNA-seq analysis. This should be fully specified and ideally controlled with isogenic wild-type and K181E clones. For transcriptional rescue, it is important to know whether the changes in wild-type organoids treated with KPT-276 are negligible. A direct wild-type comparison with or without KPT-276 is important to disentangle general drug effects from K181E-specific rescue. More detailed quantification of total TDP-43 and pTDP-43 in both nuclear and cytoplasmic fractions, including biochemical fractionation if possible, would strengthen the assertion that KPT-276 specifically reduces cytosolic pTDP-43 aggregates while sparing nuclear TDP-43.

      The organoid experiment was performed with two batches per condition to reduce the effect of batch variation. The wildtype cells and K181E mutant are derived from the same genetic background. This information is now included in the method section on page 14. Given the criticisms by review 1 and 2 on the RNAseq data, we have removed this non-essential data. 

      (16) Beyond the core issues above, several additions could greatly enhance the impact. The manuscript currently emphasizes XPO1, but the genetic and chemical data clearly implicate RNA splicing, translation, and proteostasis as equally strong or stronger regulators of TDP-43 phase states. A more integrated model that explains how these pathways intersect, for example, how splicing factor availability, ribosome loading, and proteasome capacity co-govern anisosome nucleation, growth, and hardening, would be valuable.

      We now discuss a new model in discussion based on our new Figure 6, which integrates the role of RNA splicing and nuclear transport in TDP-43 phase regulation on page 10. We agree with the reviewer that other questions are also important for future studies.

      (17) A key unresolved question is whether XPO1 is acting directly on TDP-43, or instead primarily regulates anisosomes by exporting other factors that more proximally control TDP-43 phase behavior. Given that TDP-43 is not a canonical XPO1 cargo and prior work indicates that its nuclear export is largely passive, it seems at least as plausible that XPO1 inhibition alters the nuclear concentration or localization of splicing factors, RNA-binding proteins, chaperones, or other modifiers identified in the screens, and that changes in these proteins secondarily reshape anisosome dynamics. In other words, XPO1 may be exporting a more direct regulator of anisome formation and hardening, rather than exporting TDP-43 itself in a specific, regulated way. The current data do not distinguish between these possibilities. Systematic identification of XPO1-dependent cargos that colocalize with or biochemically associate with anisosomes, combined with targeted perturbation of their nuclear export, would be needed to determine whether the relevant XPO1 substrate in this system is actually TDP-43 or an upstream modulator of its phase behavior.

      As discussed above, our new data regarding the role of RNA in TDP-43 phase regulation should alleviate this concern, although we cannot exclude the possible involvement of splicing factors in this process. We also clearly state that there is no evidence to support a direct interaction between TDP-43 and XPO1 on page 8.

      (18) Testing whether identified modifiers converge on nuclear TDP-43 concentration would be informative. Since phase separation is concentration-dependent, measuring nuclear versus cytoplasmic TDP-43 levels across key perturbations, including splicing inhibition, translation inhibition, proteasome inhibition, HSP90 inhibition, and XPO1 modulation, would help determine whether modifiers mainly work by changing nuclear TDP-43 concentration or by altering interaction networks and the material properties of condensates.

      In the newly performed immunoblotting experiment, we measured the TDP-43 levels in drug-treated cells but found no effect by most drugs (Supplemental Figure 1).

      (19) Examining other ALS-relevant RNA-binding proteins would be valuable. Given the role of XPO1 and other hits, it would be informative to briefly test whether similar principles apply to FUS, hnRNPA1, or other ALS-relevant RNA-binding proteins in the same cellular context, to argue for generality versus TDP-43-specific idiosyncrasies of the 2KQ system.

      We agree that this is an important issue but we feel the proposed experiments are beyond the scope of the study.

      (20) The Introduction sometimes implies that anisosomes are common and well-established intermediates en route to pathology. It would be helpful to more clearly state that, to date, anisosomes are primarily observed in overexpression and mutant systems and have not yet been unequivocally demonstrated in human patient tissue. The link between PDGFRβ, PAK4, GSK-3β, and YAP and TDP-43 phase dynamics is intriguing but only briefly mentioned. The authors should either expand on this or tone down the emphasis in the Results section.

      We have revised the introduction and added the following sentence on page 4. “The 2KQ-containing anisosomes, observed mostly in the nucleus under overexpression conditions, have not been validated in human patient samples.”

      (21) In the organoid methods, the authors should consider clarifying whether doxycycline is continuously used, which might alter TDP-43 expression and nuclear transport in a non-negligible way.

      The organoid model does not involve protein overexpression or doxycycline treatment. We measured endogenous p-TDP-43, which is why we feel this experiment is very significant. Unlike many other p-TDP-43 detection studies that rely on TDP-43 overexpression or exposing cells to excess stressors, we could detect substantial p-TDP-43 in 3D organoids grown under normal conditions, whereas the same cells grown and differentiated in 2D culture do not show p-TDP-43 (Zhang Q. et al., BioRxiv 2025).

      (22) For statistical methods, it would be beneficial to indicate whether multiple-comparison corrections were applied for the many FRAP, anisosome count, and size comparisons beyond DESeq2 internal corrections for RNA-seq.

      We have added more statistical information to the figure legends.

      (23) Some figure legends could more clearly indicate whether the images shown are single z-planes or maximum intensity projections and how the thresholding for anisosome detection was performed.

      We revised the figure legends to include this information. As for anisosome detection, because they are so obvious, standard thresholding combined with automated counting was sufficient to identify them.

      (24) In its current form, the manuscript contains an impressive set of screens and some nicely executed imaging of TDP-43 condensates, highlighting nuclear export among other pathways as a modulator of TDP-43 phase behavior. However, the physiological relevance is undercut by heavy reliance on an acetylation-mimetic, RNA-binding-defective TDP-43 mutant and a homozygous K181E organoid model. The mechanistic link between XPO1 and TDP-43 remains largely inferential and partly at odds with prior work. The conclusion that cytoplasmic TDP-43 aggregation is only a modest contributor to disease is not firmly supported by the available data.

      We agree with the reviewer that the strength of the study is our unbiased approach that identifies pathways capable of modulating TDP-43 phase behavior. In the revised paper, we included several experiments using an in vitro semi-permeabilized cell system to further dissect the role of nuclear export in TDP-43 phase separation. We believe that these new results should provide significant mechanistic insight that links nuclear export and RNA transcription and splicing to TDP-43 phase regulation. Additionally, we have revised our paper carefully to discuss the physiological relevance and the limitation of our study.

      (25) With substantial additional mechanistic work, particularly around XPO1, rigorous validation in more physiological TDP-43 contexts, more sensitive detection of cytoplasmic TDP-43 aggregates, and a tempering of the central claims, this study could make a meaningful contribution to understanding how nucleocytoplasmic transport and other cellular pathways influence TDP-43 phase transitions and aggregation. The work should be reframed as an important screening study that identifies nuclear export as one among several cellular processes that modulate TDP-43 phase behavior in a model system, rather than as a definitive demonstration that nuclear export governs pathological TDP-43 aggregation in disease.

      We now reframe the study as an important screening study that identifies nuclear export among several other pathways as modulators of TDP-43 phase behavior. We also propose a model that links RNA splicing to nuclear export in TDP-43 phase regulation.

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. However, significant concerns regarding experimental controls, reporting transparency, and model translatability currently limit the strength of the conclusions and the interpretability of several key findings.

      We thank the reviewer for acknowledging the significance and innovation of our study.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      Weaknesses:

      Despite its strengths, the manuscript has several major limitations that affect data interpretation and confidence in the conclusions.

      (1) Lack of appropriate controls for overexpression experiments:

      A central concern is the absence of proper controls for TDP-43 and XPO1 overexpression. Prior studies (including those cited by the authors, Archbold et al.2018) show that overexpression of WT TDP-43 alone is toxic to neurons. Thus, the experimental system itself may induce anisosome formation independently of the mechanisms under study. Similarly, XPO1 overexpression lacks a suitable control (e.g., mCherry alone or mCherry fused to a protein known to be independent of TDP-43). The near-complete colocalization of XPO1 with TDP-43 anisosomes upon overexpression raises the possibility that these structures reflect non-physiological protein accumulation rather than regulated assemblies.

      As mentioned in our response to reviewer 1, point 1, we have added more discussions to justify the use of acetylation mimetics in our study. We agree with the reviewer that these large puncta (both anisosomes and gel-like structures) likely resulted from TDP-43 overexpression. Nevertheless, in a titration experiment done by Yu et al. 2020 (PMID: 33335017), they showed that ectopic TDP-43 undergo demixing even at concentrations lower than endogenous TDP-43, although the demixed puncta were very small. Their result suggested that overexpression per se does not change TDP-43 phase behavior, only enlarge the demixed TDP-43 structures, which is necessary for our screen and imaging-based characterization.

      For XPO1 overexpression, we have done the mCherry alone control but due to space limit in Figure 5, we did not include it. We now include the data in Supplemental Figure 4. This figure shows that overexpression of mCherry did not change TDP-43 localization or anisosome structures.

      (2) Insufficient experimental and analytical transparency:

      The manuscript frequently lacks clear reporting of experimental details. In multiple figures, the stated number of independent experiments does not match the number of data points shown, making it difficult to assess statistical validity. Concentrations used in the compound screen are not clearly defined, nor is it stated whether multiple concentrations were tested. It is unclear how many wells, cells, or independent cultures were analyzed. The criteria used to reduce 1,533 screening hits to 211 candidates via STRING analysis are not explained. Knockdown and overexpression efficiencies are not reported.

      We apologize for these omissions. We have added more experimental details to the figure legends and the method. For the imaging experiments, data points reflect randomly selected individual cells imaged in 2-3 independent biological repeats. This is now stated in the figure legends. For chemical screens, we screened against NCATS libraries was first done at top concentration (10 mM) to ensure inhibitory efficacy for all potential hits. In the follow-up validation study, we validated the top hits using a series of concentrations, as shown in Figure 1B. Drug concentrations are provided in Figure 2A, 4A, C, E, F, 5A-D, F, Figure 6F, G, Figure 7A)

      We explain the STRING analysis in more detail now. Basically, STRING is a protein-protein interaction network that reports all potential interactions between any proteins in human proteome. Given the potential off-target effect of siRNA, we assume that if the screen identifies multiple components of a protein interaction network or pathway, the result is more likely to be real.

      We did not check XPO1 knockdown efficiency in high through-put screens (HTS) for several reasons. Firstly, the large number of positive hits makes it impossible to check knockdown efficiency for all of them. Secondly, the effect of XPO1 knockdown on anisosomes was seen with 6 different siRNAs in two rounds of screens. Thirdly, in the HTS protocol, we routinely included a transfection control (siRNAdeath) to control transfection efficiency. We would only process the data if siRNAdeath control killed > 90% of the cells. Lastly, the XPO1 knockdown result was independently validated by small molecule inhibitors. For TDP-43 overexpression, the study by Yu and colleagues suggested that the expression is more than 20-fold higher than endogenous TDP-43, but they showed that anisosome formation is not an artifact of protein overexpression. When the expression level was titrated down, they could still detect anisosomes.

      (3) RNA-seq concerns:

      The RNA-seq experiments are particularly problematic. The number of biological replicates per condition is not stated, and heatmaps suggest that only one sample per group may have been used, which would preclude statistical analysis. No baseline comparison between WT and mutant TDP-43 is shown. Given that TDP-43 is an RNA-binding protein, splicing analyses would be far more informative than gene expression alone, yet no splicing data are presented. Moreover, nuclear retention of TDP-43 does not preclude nuclear aggregation, which may still impair its splicing function.

      We apologize for the lack of clarity regarding the RNA-seq design. For each condition, organoids of two independently differentiated batches were treated in triplicate. What we showed before was averaged expression levels. We pooled the organoids of the same treatment from the two batches to reduce the impact of batch variation.

      Given the criticisms from both reviewers 1 and 2 on the limited interpretation power of the RNAseq study, we have removed this data from the revised manuscript.

      (4) Limited translatability to neuronal biology:

      All anisosome analyses are performed in a cancer cell line, raising concerns about relevance to post-mitotic neurons. While organoids are used as a secondary model, the assays performed do not overlap with those used in cancer cells, making it difficult to assess whether anisosome-related mechanisms are conserved. Neuronal toxicity, a critical outcome given known TDP-43 biology, is not assessed. Prior work has shown that WT TDP-43 overexpression alone is toxic to neurons, yet this is not addressed.

      We agree with the reviewer that the model used in this study is not directly relevant to neurodegeneration. However, as pointed out by the reviewer, neurons are much more sensitive to TDP-43-associated toxicity. By contrast, the cell line used in this study can tolerate TDP-43 overexpression with no detectable cytotoxicity. This feature makes it feasible to evaluate how different cellular processes modulate TDP-43 phase behavior without the confounding effect from cytotoxicity. Notably, the processes identified by our screens are all house-keeping pathways that are conserved in neurons. Thus, we believe that the reported findings are likely applicable to neurons. That being said, we have revised our paper to ensure that we don’t overstate the clinical relevance of our work.

      (5) Conceptual and interpretational gaps:

      The authors quantify anisosome number but also report conditions in which anisosome number decreases while size increases. The biological interpretation of larger anisosomes is not discussed, and whether this reflects improvement or worsening of pathology is unclear. Compounds targeting the same mechanism (e.g., nuclear export inhibition) are inconsistently used across experiments (KPT compounds, verdinexor, leptomycin B), raising concerns about reproducibility. In organoids, the experimental paradigm shifts to long-term treatment (35 days vs. 16 hours), further complicating interpretation.

      We thank the reviewer for these critical points. As pointed out by the reviewer 1 in point 4 above, we do not have evidence to establish a convincing correlation between the size of anisosomes and clinical phenotypes. Regarding the use of different drugs for different experiments, the initial screen identified KPT and Verdinexor because they are investigational drugs, but Leptomycin B was not in our library. In the follow-up studies, we switched to Leptomycin B because 1) it is highly potent and specific; 2) it was better characterized and more commonly used as inhibitors of XPO1 according to the literature. However, for the organoid study, we had to switch back to KPT because of the toxicity issue associated with long-term application of Leptomycin B.

      (6) Overinterpretation of rescue effects:

      Although the authors state that they aim to test whether nuclear export inhibition rescues neuronal defects, no functional neuronal readouts are provided (e.g., viability, morphology, axon outgrowth, or electrophysiological measures). RNA-seq alone is insufficient to support claims of rescue.

      Our interpretation of the RNA-seq data was that the rescue effect by nuclear export inhibition was limited and probably insignificant. Given that this negative data is not conclusive, we have removed it from the revised manuscript.

      (7) Finally, the model does not appear to exhibit cytosolic TDP-43 aggregation at baseline. It remains unclear whether longer induction would produce cytosolic gel-like assemblies and whether these would be prevented by nuclear export inhibition. Long-term data are shown only in organoids, yet anisosome formation is not assessed there.

      The expression system used in the study reaches a steady state after 24 h of induction. Prolonged expression up to 48 h did not alter the number of anisosome, nor does it change TDP-43 phase behavior. We now clarify this point on page 4.

      Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      We thank the reviewer for acknowledging the significance and strength of our study.

      Weaknesses:

      The mechanisms underlying the connection between nuclear export and phase transition need further clarification. Broader consequences of XPO1 inhibition are not addressed.

      We agree that our previous manuscript did not address how nuclear export inhibition affect TDP-43 phase behavior. As discussed in our paper, we proposed that the effect of nuclear export inhibition on TDP-43 phase separation is likely indirect. The most likely scenario is that inhibition of nuclear export changes the nuclear environment over time, which affects TDP-43 phase separation. We have tried to isolate nuclear extracts from control and LMB-treated cells and used mass spectrometry to identify proteins that are differentially present in the nucleus. However, knockdown of the identified top candidates did not abolish LMB-induced phase alteration (not shown). Considering our observation that RNA splicing is another modulator of TDP-43 phase behavior, we reasoned that it is possible that it is the combined change of RNA and protein composition in the nucleus that alters TDP-43 phase behavior. In new experiments presented in Figure 6, we now used a semi-permeabilized in vitro system to demonstrate that LMB treatment stabilized anisosomes in an RNA-dependent manner (see response to point 4 by reviewer 1). This new data allows us to propose a new model that link RNA splicing and nuclear export in TDP-43 phase regulation (Discussion).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Include appropriate controls for all overexpression experiments. In particular, overexpression of WT TDP-43 alone and suitable tag-only controls (e.g., mCherry alone or mCherry fused to a protein unrelated to TDP-43/XPO1) should be included to control for aggregation driven by non-physiological protein levels.

      In Supplemental Figure S4, we included a tag-only control, which shows that mCherry alone does not affect the localization of XPO1, neither did we see mCherry co-localizes with TDP-43.

      Since WT TDP-43 itself does not form anisosome and because the goal of the study was to test how anisosome dynamics is affected by various conditions, we did not repeat our experiments with WT TDP-43.

      (2) Address whether TDP-43 anisosomes form under endogenous or near-physiological expression levels. If possible, include experiments using lower expression systems or endogenous tagging to demonstrate that anisosome formation is not solely an overexpression artifact.

      As mentioned above, in a titration experiment done by Yu et al. 2020 (PMID: 33335017), they showed that ectopic TDP-43 undergoes demixing even at concentrations lower than endogenous TDP-43, although the demixed puncta are small. Their result suggested that overexpression per se does not change TDP-43 phase behavior. Instead, it only enlarges the demixed TDP-43 structures, which is necessary for our screen and imaging-based characterization.

      (3) Clearly define biological versus technical replicates throughout the manuscript and report exact n-numbers for all experiments in figure legends and/or methods. Resolve discrepancies between stated and displayed n-numbers (e.g., figures showing more data points than the number of independent experiments reported). Further, include how data points were defined (e.g., cells, fields of view, wells).

      We now state clearly the biological repeats in figure legends. We did not use N number to specify technical replicate. The discrepancy between the stated N number (biological repeats) and the data points is because for imaging experiments, data points usually represent single cells collected from 2-3 biological replicates (N=2 or 3). Data points are now clearly defined in the figure legends (anisosome, cell, imaging field, or independent experiment).

      (4) The authors state that they identified a list of compounds that reduced anisosomes. Please clarify how the threshold was determined: Was this a statistical analysis or a specific threshold that has been used?

      For both siRNA screen and chemical genetic screen, we calculated the Z-score and used Z-score>2 as a cutoff. This is mentioned in the method.

      (5) Provide a complete list of compounds used in the chemical screen, including concentrations tested and whether multiple doses were evaluated.

      As mentioned above, the initial screen was done with just one concentration (10 mM). Identified positive hits were re-tested with multiple doses as shown in Figure 1. The compounds are from a commercial library (LOPAC R1280, Sigma #LO4200). The list of compounds can be found at vender’s website.

      (6) Clearly explain the criteria used to reduce the initial 1,533 screening hits to 211 candidates following STRING analysis, including cutoffs and prioritization logic.

      We now explain that the Z-score was used to further narrow down the hit (page 6). Additionally, we provide an explanation on how we use STRING to further narrow down the list. The sentence reads as “To further narrow down the list, we performed a STRING protein network analysis based on the assumption that a protein interaction network bearing multiple positive hits would be more likely to be a true effector.”

      (7) Report knockdown and overexpression efficiencies for all genetic perturbations used in the study.

      For TDP-43 overexpression, the study by Yu and colleagues suggested that the stable cell line expresses 20-fold more TDP-43 than endogenous one, but they showed that anisosome formation is not an artifact of protein overexpression. When the expression level was titrated down, they could still detect anisosomes (Yu, H. et al., Science 2021). For knockdown efficiency, since the screen used 6 different siRNAs for each identified target (a few hundred), it is technically challenging to validate the knockdown efficiency of each siRNA by conventional qRT-PCR. To control knockdown efficiency, we transfected cells in parallel with siRNA-death that contains a mixture of siRNAs targeting several essential genes (Qiangen, #1027299). We would only process the data if siRNAdeath control killed > 90% of the cells, indicating good knockdown efficiency.

      (8) Clarify the biological interpretation of changes in anisosome size versus number, particularly in conditions where fewer but larger anisosomes are observed. Discuss whether larger assemblies are hypothesized to be protective, neutral, or deleterious.

      Live cell imaging was used to dissect why cells treated with certain drugs such as XPO1 inhibitors have fewer but larger anisosome. Figure 5F shows that this is caused by the fusion of small anisosomes. Our data does not suggest that the size of anisosomes can differentiate between protective or deleterious state, but rather it is the LLPS state and subcellular localization of these assemblies that may play a more critical role in determining whether TDP-43 forms deleterious protein aggregates. The discussion is on page 10.

      (9) Specify whether all anisosomes induced by XPO1 overexpression were gel-like or whether this applied only to a subset. If only a subset was affected, please provide quantifications, otherwise state clearly that all anisosomes in XPO1 overexpression were gel-like.

      All TDP-43 puncta mislocalized to the cytoplasm in XPO1-overexpressing cells are gel-like because the FRAP experiment in Figure 5I was done with randomly selected TDP-43 puncta mislocalized to the cytoplasm.

      (10) Clarify which anisosomes (nuclear vs cytosolic; gel-like vs non-gel-like) were selected for FRAP analyses in Figure 5I.

      For Figure 5I, the control anisosomes in untreated cells are nuclear while under mCh-XPO1 expressing condition, only those in the cytoplasm were randomly selected for photobleaching.

      (11) The translatability of the conclusion based on cancer cell lines to brain organoids is not convincingly shown and could be strengthened by including additional assessment of anisosomes. While this might not be feasible in 3D cultures, the authors could alternatively use 2D cultured neurons to perform the same assays as performed in the cancer cell line. Additionally, the same treatment strategy should be applied. The reasoning for increasing treatment to 35 days in the organoids is unclear.

      In another manuscript that is currently under revision, we compared 2D iNeuron culture with 3D organoids. A pre-print is available at https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full. In this study, we found that endogenous TDP-43 K181E mutant do not undergo phosphorylation-dependent transition to aggregate in 2D cultures. Only when these cells were grown into 3-D organoids, TDP-43 phosphorylation could be detected. (see supplemental Fig. S1c, d in https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full). Thus, it is not possible to repeat the experiments in this study in 2D iNeuron cultures. We agree with the review that there is a gap between the study using the cancer cell line and the use of K181E iPSC-derived 3D organoids. We have toned down our conclusions throughout the text.

      (12) Address neuronal vulnerability explicitly by assessing toxicity, viability, or functional neuronal readouts, particularly given prior reports that WT TDP-43 overexpression alone is neurotoxic.

      We agree that this is an important point, but the main goal of this study was to dissect the cellular pathways/mechanisms that govern TDP-43 phase separation. We feel that the requested experiments are beyond the scope of the current study.

      (13) Clearly state the number of biological replicates used for each RNA-seq condition. Establish baseline transcriptional differences between WT and mutant TDP-43 prior to assessing the effects of nuclear export inhibition. Include PCA plots and heatmaps, including all samples.

      As mentioned above, we have decided to remove the RNAseq data from the manuscript to save room for new results.

      (14) Given the role of TDP-43 as an RNA-binding protein, consider including splicing analyses to assess whether nuclear export inhibition preserves or disrupts TDP-43-dependent RNA processing.

      We thank the reviewer for this suggestion. However, we feel that the proposed experiments are beyond the scope of the current study.

      (15) Improve clarity of transcriptomic visualizations (e.g., GO-term plots) and explicitly define all group labels used (e.g., Group A vs Group B).

      We have removed the RNAseq data.

      (16) Ensure consistent use of disease terminology (ALS vs FTD) throughout the manuscript, e.g., lines 222 and 244.

      We have checked the usage of these terms to make sure they are accurately used.

      (17) Correct figure and axis labeling errors (e.g., Figure 3A x-axis range).

      Figure 3A indicates the Z score distribution of the entire human genome. As stated on page 6, 21,404 genes were targeted.

      (18) Avoid overstatements in the Discussion that are not directly supported by the presented data, particularly regarding the interpretation of proteasome inhibition and gel-like anisosome states.

      We have revised our discussion substantially to tone down our conclusions.

      (19) Clarify the rationale for switching between different nuclear export inhibitors across experiments and discuss whether results were consistent across compounds.

      In the acute experiments down with the cancer cell line, we used LMB because it is potent and well characterized. In organoid experiment, we switched to KPT-276 because it is better tolerated by organoids, especially during longer treatment.

      Reviewer #3 (Recommendations for the authors):

      Major concerns that require clarification or further strengthening:

      (1) The connection between nuclear export and liquid-solid phase transition is not clear. The 2KQ mutant forms nuclear anisosomes. The manuscript does not provide data about its nuclear-cytoplasmic distribution normally, nor how the distribution is changed upon nuclear export inhibition or enhancement. In Figure 5I, it is unclear whether the anisosomes are in the nucleus or cytoplasm. The dynamics of nuclear vs cytoplasmic anisosomes should be measured separately. What is the mechanism that promotes nuclear export and changes the dynamics, especially nuclear anisosomes?

      As mentioned by the reviewer, the 2KQ mutant forms anisosomes only in the nucleus. This was documented in Yu, H. et al., Science 371 (2021), and also shown in our Figure 4A, F, Figure 5A. Figure 5A also shows that nuclear export inhibition does not change anisosome localization, only making them bigger while reducing the numbers. For Figure 5I, the control anisosomes in untreated cells are nuclear while under mCh-XPO1 expressing condition, only those present in the cytoplasm were randomly selected for bleaching.

      (2) Figure 5J, no obvious XPO1 is sequestered to anisosomes, as described in lines 208-209.

      Unlike Figure 5G, this experiment studied the localization of endogenous XPO-1 by immunostaining. As discussed in Yu et al., Science 371 (2021), proteins inside anisosomes could not be stained by antibodies due to an accessibility problem. This explains why we could only detect reduced XPO1 after anisosome induction.

      (3) Figure 6A, the localization of phosphor-TDP-43 is not clear. And it is not clear what cell types contain the aggregates. Higher-resolution images need to be included. The mechanism by which XPO1 inhibition reduces TDP-43 aggregation requires further validation. It remains unclear whether it is directly mediated through altered nucleocytoplasmic transport of TDP-43.

      We agree that it is technically challenging to visualize the precise subcellular localization of p-TDP-43 in 3D organoids. In the manuscript that reports the characterization of the 3D organoids, we dissociated cells from the 3D organoids by trypsin digestion and plated them out in 2D before immunostaining and imaging. We could clearly see p-TDP-43 co-localizes with the neuronal marker TUJ1 and is localized outside of nucleus (see figure 1 of https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full)

      In the newly added Figure 6, we used a semi-permeabilized cell system to dissect the phase separation dynamics of TDP-43 2KQ in cells treated with the nuclear export inhibitor LMB. Our data suggests that nuclear export inhibition alters the nuclear environment, making it more favorable for the liquid phase of TDP-43. This is dependent on nuclear RNA.

      (4) XPO1 controls the export of numerous essential proteins, and its inhibition can produce broad, potentially toxic effects unrelated to TDP-43. The manuscript should include a discussion of these off-target consequences.

      We thank the reviewer for this point. Given the new data in Figure 6, we now add some more discussion on the potential mechanism by which nuclear export inhibition modulates TDP-43 phase separation. This can be found on page 10.

      References:

      Zhang, Q. et al. A human forebrain organoid model phenocopies dysregulated RNA and protein homeostasis in ALS/FTD-associated TDP-43 proteinopathies. bioRxiv (2025). (https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full

    1. eLife Assessment

      This study presents useful findings on the molecular mechanisms driving female-to-male sex reversal in the ricefield eel (Monopterus albus) during aging, which would be of interest to biologists studying sex determination. The manuscript describes an interesting mechanism potentially underlying sex differentiation in M. albus. However, the current data are incomplete and would benefit from more rigorous experimental approaches for Western blotting.

    2. Reviewer #1 (Public review):

      Summary:

      This preprint investigates the molecular mechanism by which warm temperature induces female-to-male sex reversal in the ricefield eel (Monopterus albus), a protogynous hermaphroditic fish of significant aquacultural value in China. The study identifies Trpv4 - a temperature-sensitive Ca²⁺ channel - as a putative thermosensor linking environmental temperature to sex determination. The authors propose that Trpv4 causes Ca²⁺ influx, leading to activation of Stat3 (pStat3). pStat3 then transcriptionally upregulates the histone demethylase Kdm6b (aka Jmjd3), leading to increased dmrt1 gene expression and ovo-testes development. This work aims to bridge ecological cues with molecular and epigenetic regulators of sex change and has potential implications for sex control in aquaculture.

      This revision is an improvement to the manuscript. However, there are still several remaining issues that are not resolved and that limit enthusiasm.

      (1) The Supplementary File 1 contains a compilation of Western blots. However, the control protein (for example GAPDH) is on a *different gel* in all of the tabs. For best practices, the protein that is used as the "loading control" needs to be on the same membrane (same Western blot), not on a different blot. It is not compelling to normalize a loading control protein on a separate blot. This reduces enthusiasm for all of the protein data in the manuscript.<br /> a. The blots under the tab "Fig. 5D" are dirty and the blot the GAPDH is over-exposed.

      (2) The images provided in the response to authors have no legends and are not explained in the text. As such, they are not supportive data in their current form.

      (3) The antibodies that were listed as "home-made" need to be described in great details. For example, we need to know the species that the antibodies were generated in. Additionally, we need to know the antigen (amino acid residues of the recombinant protein).

      (4) The reference genes for the qRT-PCR are not listed in the Materials and Methods. The authors need to list the reference gene and tell us why they selected those genes.

      (5) The comparison of the turtle and ricefield eel of kdm6b should be shown as a supplementary file and not listed as data not shown.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      This preprint investigates the molecular mechanism by which warm temperature induces female-to-male sex reversal in the ricefield eel (Monopterus albus), a protogynous hermaphroditic fish of significant aquacultural value in China. The study identifies Trpv4 - a temperature-sensitive Ca²⁺ channel - as a putative thermosensor linking environmental temperature to sex determination. The authors propose that Trpv4 causes Ca²⁺influx, leading to activation of Stat3 (pStat3). pStat3 then transcriptionally upregulates the histone demethylase Kdm6b (aka Jmjd3), leading to increased dmrt1 gene expression and ovo-testes development. This work aims to bridge ecological cues with molecular and epigenetic regulators of sex change and has potential implications for sex control in aquaculture.

      Strengths:

      (1) This study proposes the first mechanistic pathway linking thermal cues to natural sex reversal in adult ricefield eel, extending the temperature-dependent sex determination paradigm beyond embryonic reptiles and saltwater fish

      (2) The findings could have applications for aquaculture, where skewed sex ratios apparently limit breeding efficiency

      Weaknesses:

      Although the revised manuscript represents an improvement over the original version, substantial weaknesses remain.

      We thank you for the critical comments. We have responded to your concerns by a point by point manner, and please see detail below.

      Scientific Concerns

      (1) Western blot normalization and exposure: The loading controls (GAPDH) in Fig. S3C appear overexposed, as do several Foxl2 blots. Because these signals are likely outside the linear range, I am not convinced that normalization is reliable. This raises concerns about the validity of the quantified results.

      We thank you for the concerns. We have repeated the experiments, and new blots were loaded in Fig.S3C.

      (2) Antibody validation and referencing (Line 776): The authors need to refer explicitly to figures demonstrating antibody validation. At present, these data are provided only as a supplementary file that is not cited in the manuscript. In addition, the Sox9a antibody appears to yield indistinguishable signals in control and RNAi conditions, suggesting that it may not recognize eel Sox9a. This issue is not addressed by the authors. Furthermore, antibody validation Western blots should be quantified.

      We thank you for the comments. We have repeated the siRNA experiments to show the specificity of the antibodies used. This file, named as the supplementary file 1, is now cited in “WB analysis” in the Materials and Method part. As required, the antibody validation of WB are uploaded in the supplementary file 1. Antibody validation for WB are now quantified, and please see the new figure 3 and supplementary Figure 3.

      (3) Unclear sample sizes (N values): Sample sizes remain unclear for several figures:

      (a) Fig. 3F - No N value is provided. Each graph shows three data points; does this indicate that only three samples were quantified? If ten samples were collected, why were all not quantified?

      We apologize for the confusion. Three data points were previously used to shown data of 3 replicates. In new figure 3F, 10 randomly selected sections were imaged, and the data are shown. In the revised manuscript, the sample numbers (the N values) are added, and all the information can be found in the figure legend.

      (b) Fig. 4 - No N values are reported.

      Now N values are added. Please see the figure legend.

      (c) Fig. 5A - Again, only three data points are shown per group, despite the apparent availability of twelve samples. The rationale for this discrepancy is not explained.

      We apologize for the wrong data representation. Now all the data points are shown in Figure 5.

      (4) qRT-PCR normalization: The manuscript does not specify the reference gene(s) used for qRT-PCR normalization. Although expression levels are reported as "relative," neither the identity of the reference gene(s) nor the justification for their selection is provided.

      We now have specify the reference gene in “Quantitative real-time PCR (qPCR) experiments” part in the Materials and Methods section.

      (5) Specificity of key antibodies: While the authors have made some effort to validate anti-Amh, anti-Sox9, and anti-Dmrt antibodies, the results remain incomplete. The Amh and Dmrt antibodies detect reduced protein levels following knockdown of their respective targets, which is encouraging. However, the Sox9a antibody shows no difference between control and RNAi conditions, suggesting it does not recognize eel Sox9. This is not acknowledged in the manuscript. In addition, no validation data are presented for Foxl2. Antibody validation data must be clearly referenced in the main text and presented in an interpretable and quantitative manner.

      The antibody specificity is very important. For that reason, we have generated at least two different antibodies for each target protein, using full-length or small peptide as antigen. We have repeated the experiments for key antibodies such as Dmrt1 and Sox9a. IF and WB results clearly showed the specificity of the antibodies.

      Author response image 1.

      Foxl2 antibody has also been reported in ricefield eel (Hu et al. SCIENTIFIC REPORTS | 4: 6884 | DOI: 10.1038/srep06884, Molecular cloning and analysis of gonadal expression of Foxl2 in the ricefield eel Monopterus albus).

      After short term warm temperature exposure, only a small portion of somatic cells in ovary may be induced to express the male markers. As different techniques have different capacity (sensitivity), some techniques were more easy to detect that change. For instance, qPCR and WB are ready to detect it, whereas IF is a little difficult in obtaining good quality data.

      (6) Immunofluorescence data quality: The immunofluorescence images remain difficult to interpret. I strongly encourage the authors to enlarge the image panels and to present monochrome images (white signal on black background). The current presentation severely limits interpretability.

      We thank you for the comments. We think that our IF images are of decent quality. Due to the limits of the Figure space (already busy for Figure 3), enlarging the image panels or presenting additional monochrome images will compromise the quality of other data. Alternatively, if you still concern its quality, we can put it in the supplementary.

      Author response image 2.

      (7) Unreferenced supplementary figure: Fig. S4 is included in the submission but is not referenced anywhere in the manuscript text.

      We now have renamed the supplementary Figures. And we have double checked the text to make sure all Figure information is correctly referenced. Figure S4 is removed, as it is not necessary.

      (8) Fig. 5B image resolution: The micrographs in Fig. 5B are too small to allow meaningful evaluation of the data.

      Now new Figure 5B images with higher resolution were shown.

      (9) Unexplained data inclusion (Fig. 5E): Fig. 5E includes a pERK blot that is not mentioned in the Results section. The rationale for including these data is unclear.

      Previous work have shown that FGF/ERK signaling may play a role in sex change of ricefield eel (in Chinese). We therefore examined the Erk activity to explore whether it is involved in sex reversal. The results showed that pErk was comparable between ovary and ovotestis. At your suggestion, we decided to remove the data.

      (10) Poor blot quality (Fig. S3C): The blots in Fig. S3C exhibit high background and overexposure. I am concerned about the reliability of the quantification shown in panel D.

      The experiments have been repeated at least three times, and similar results were obtained. We now have replaced some of the WB that were of high background or overexposure.

      (11) Poor blot quality (Fig. S5G): The Stat3 blots in Fig. S5G contain numerous white artifacts, raising concerns about their suitability for normalization in panel H.<br />

      We now have repeated the experiments, and uploaded a new representative blot with better quality.

      (12) Missing controls (Fig. 6E): Fig. 6E lacks controls for HO-3867 and Colivelin treatments alone. Without these controls, it is not possible to determine whether the reported effects are meaningful.

      We thank you for the comments. We now have added the data required (with HO-3867 and Colivelin treatments alone).

      (13) Graphical presentation: The use of a light blue-to-pink gradient in bar graphs throughout the manuscript does not aid interpretation. I recommend using more distinct colors (e.g., red, orange, green, blue, purple, gray, black) to improve clarity.

      We thank you for the comments. We now have changed the blue-to-pink gradient to more distinct color system to better present the data. Please see the detail in the revised Figures.

      In summary, the interpretation of the study remains limited by persistent issues related to data presentation, image quality, and reagent specificity.

      We thank you for the critical comments about our data, in particular for antibody specificity and image quality, and the detailed instruction for how to better present the data. Answering your questions have greatly improved the quality of the manuscript. We admit that due to the technique challenging (with different conditions and different doses of small molecules) and higher cost of animal experiments, some of the WB or IF experiments may not be of high standards.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Editorial Concerns

      (1) Overstatement of conclusions: In lines 16-18, the authors state that Trpv4 "mediates" warm temperature-driven sex reversal. This claim is too strong given the data and should be toned down.

      We agree with our editorial comment about the overstatement. Now it reads “Trpv4 links environmental temperature to testicular differentiation in ricefield eel”.

      (2) Misuse of statistical language (Line 213): The term "significant" is used where statistical significance was not measured. The wording should be revised.

      We thank you for the point, and now have replaced “significant” to “marked”.

      (3) Terminology (Line 238): The term "co-expression" is inaccurate in this context. I suggest replacing it with "co-upregulation."

      We thank you for the point, and have changed it accordingly.

      (4) Drug description errors (Lines 241-242): The manuscript incorrectly identifies which drug functions as an agonist and which as an antagonist. This caused considerable confusion and must be corrected.

      We have carefully checked the sentence, and it was correct, as RN1734 and GSK1016790A are known Trpv4 specific antagonist and agonist, respectively.

      (5) Gene examples missing (Lines 247-250): The authors should explicitly name the testis-biased and ovary-biased genes referred to in this section.

      We thank you for the point, and now it reads “warm temperature exposure increased the expression of testicular differentiation genes such as dmrt1 and gsdf, accompanied by moderately decreased expression of ovarian differentiation genes such as cyp19a1a and foxl2”.

      (6) Lack of experimental context (Lines 322-324): Rather than simply listing the drugs used, the authors should briefly explain what each compound inhibits or activates and why it was employed.

      We have described this in the manuscript. The information of pStat3 activator and inhibitor has been described in Lines 305-309, as “HO-3867, a curcumin analogue, is a selective pStat3 inhibitor, which blocks pStat3 activity by directly binding to Stat3 DNA binding domain, and Colivelin is a potent synthetic peptide activator of pStat3, which increases pStat3 levels by acting through the GP130/IL6ST complex”, and the rationale has been stated in lines 32--322 as “To functionally demonstrate that pStat3 signaling is downstream of Trpv4, rescue experiments were performed by injecting into ovaries with individual and combined small molecules”.

      (7) Discussion of evolutionary differences: The Discussion misses an important opportunity to address why Stat3 activates kdm6b in ricefield eel but represses it in turtles. It is difficult to reconcile how the same transcription factor could exert opposite effects on the same gene during sex determination without additional context. A comparison of kdm6b regulation and sequence conservation between turtles and ricefield eel would strengthen this section.

      We have downloaded the promoter sequences of red eared turtle and ricefield eel. Based on the DNA sequences (Author response image 3), the similarity (conservation) was low between the two species.

      Author response image 3.

      It was appeared that DNA around the Stat3 binding sites in turtle are GC rich (CpG island), which may be subjected to DNA methylation modification, whereas the DNA in ricefield eel are not GC rich.The observations imply that the role of pStat3 is to promote the repression of kdm6b in turtle but the activation of kdm6b in ricefield eel.

      Moreover, our unpublished data showed that Trpv4-controlled calcium signaling is required to remove the repressive histone modification H3K27me3 at the kdm6b gene. If pStat3 is downstream of Trpv4 in this case, it supports again that Trpv4-pStat3 axis activate kdm6b in ricefield eel.

      Warm temperature promotes female sex in turtle but male sex in ricefield eel. If pStat3 is mediating Trpv4, it is not surprising that it represses kdm6b in turtle but activate it in ricefield eel.

      Based on above, we have added some sentences in the discussion part, and it reads “We reasoned that a yet-unidentified co-factor may determine whether Stat3 is a transcriptional repressor or activator. A comparison of promoter sequences of kdm6b between turtle and ricefield eel supported this”.

      (8) Supplementary figure formatting: Supplementary figures should be provided in accordance with eLife formatting guidelines.

      We have now formatted the supplementary figures that are in accordance with eLife formatting requirement. Please see the new uploaded supplementary figures.

      In sum, the interpretations are still limited by the above concerns regarding data presentation and reagent specificity.

      We thank our editor for the inspiring comments. We believe we have addressed all the major concerns by our editor.

    1. eLife Assessment

      This study provides a valuable advance in understanding how disordered proteins interact with cell membranes by identifying the sequence rules that enable aromatic residues to penetrate deeply into the membrane interior. The integration of complementary computational approaches, including molecular simulations, large-scale sequence analysis, and the development of an online prediction server, makes the work potentially impactful for the membrane protein and intrinsically disordered protein communities. The evidence supporting the main conclusions is generally convincing, although its transferability across diverse membrane compositions and its validity as a prediction tool for real protein-membrane systems remain to be further established.

    2. Reviewer #1 (Public review):

      Summary:

      This work investigates the membrane insertion of aromatic-centered sequences in IDPs. Using a combination of all-atom MD simulations, the PPM method, and development of the sequence-based predictor AroMIP, the authors aim to establish a quantitative membrane insertion role for aromatic-centered motifs. The study demonstrates that flanking aliphatic and basic residues promote membrane insertion, whereas acidic and polar residues suppress insertion, and further reveals a difference between F/W-centered motifs and Y-centered motifs. The resulting AroMIP model achieves high predictive accuracy on human IDPs and is implemented as a publicly accessible web server.

      Strengths:

      This work addresses an important biological problem, as aromatic-driven membrane insertion remains poorly characterized despite mediating diverse functions like membrane remodeling and signaling. A key strength is the combination of complementary approaches, e.g., MD simulations provide mechanistic insight into insertion pathways, while PPM enables exhaustive sequence space exploration. The large-scale analysis clearly establishes L and R as promoters and E, N, and G as suppressors. The work also provides valuable mechanistic insight into how aromatic, aliphatic, and basic residues cooperate to stabilize membrane insertion states. Another important strength is the development of AroMIP as a practical prediction tool with a user-friendly online server that appears computationally efficient and broadly accessible to the community. The work is also well connected to prior experimental and computational literature, and the authors carefully position their findings within existing knowledge of membrane-associated IDPs.

      Weaknesses:

      A primary limitation is the heavy reliance on computational modeling. Training for AroMIP is generated using PPM rather than direct experimental measurements, and so the model may primarily reproduce PPM behavior rather than true membrane insertion thermodynamics. Moreover, all simulations use a single lipid composition (POPC:POPS:PIP₂ 70:25:5), but biological membranes vary substantially in cholesterol, cardiolipin, and acidic lipid content. Whether AroMIP's predictions transfer to diverse lipid environments remains untested. The 5% PIP₂ concentration used in the simulations is higher than that of a normal mammalian cell and may therefore overemphasize electrostatic contributions. Applicability beyond short 9-residue motifs is unclear, as longer-range interactions or secondary structure in full-length IDRs could modulate insertion in ways the current model does not capture. This could be considered for future development.

    3. Reviewer #2 (Public review):

      Summary:

      The paper addresses an interesting problem. The authors develop a method to assess the probability of insertion of aromatic residues in intrinsically disordered regions of proteins, to insert in the interfacial regions of membranes.

      Strengths:

      (1) The idea of the article seems very interesting. The problem of membrane association mediated by aromatic residues is definitely worth studying. Aromatic residues, especially Tryptophan (W), but also, albeit to a lesser extent, Phenylalanine (F), and Tyrosine (Y), are well known to partition preferentially to the headgroup region of the lipid bilayer.

      (2) The authors propose to decipher the sequence code for insertion of sequences containing aromatic residues in the membrane employing three types of calculation methods with decreasing order of detail and complexity, but increasing order of efficiency. First, all-atom MD simulations; second, the PPM method (protein positioning in membranes) from Lomize et al (2006), Protein Sci 15, 1318; and third, AroMIP, a mathematical model developed by the authors. The results obtained with the different simulations and mathematical methods are internally consistent.

      Weaknesses:

      (1) Aromatic residues have been shown to partition preferentially to the headgroup region of the lipid bilayer. Most of the papers on this problem were published in the mid 1990s to early 2000s. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev. Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602-608; Killian & von Heijne, TIBS 2000, 25, 429-434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, none of these articles is cited.

      (2) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth"."

      (3) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is incorrect even in a qualitative sense.

    4. Reviewer #3 (Public review):

      Summary:

      This is a well-written manuscript that describes three robust and complementary computational approaches to unravel the sequence determinants of membrane insertion, specifically of intrinsically disordered regions (IDRs) containing aromatic-centered insertion motifs.

      Strengths:

      A robust, multifaceted computational approach employing aromatic-centered model membrane-insertion peptides, which provides critical insights into the determinants of membrane insertion.

      Weaknesses:

      I only have specific concerns about some of the models used for this purpose.

      (1) Membrane composition and lipid shape characteristics: The authors chose to use a model membrane bilayer of a distinct lipid composition, POPC: POPS: PI4,5P2 (70:25:5 molar ratio), for their all-atom simulations of the various model peptides. While this may be pertinent for some of these peptides, it is not for many, such as sequence 2 derived from Drp1, which preferentially binds target conical lipids such as cardiolipin (CL) and phosphatidic acid (PA). The rationale behind using PI4,5P2, which can induce positive membrane curvature when sequestered, versus CL and PA, which both induce negative membrane curvature, is not explained.

      (2) Parallel vs. perpendicular peptide orientation of sequence 2 in peripheral Drp1-lipid interactions: On page 11, the authors state that their simulation results of sequence 2 derived from Drp1 "contrasts with a transmembrane orientation proposed by Mahajan et al." However, upon review, a transmembrane orientation for this region has never been proposed anywhere. Drp1 is a peripheral membrane protein that reversibly binds CL- and PA-containing membranes via its intrinsically disordered variable domain containing an aromatic-centered WRG motif. Indeed, the model presented in Figure 9 of Mahajan et al. displays a peripheral and parallel orientation of the transiently helical WRG-containing motif rather than a transmembrane (i.e., across the bilayer) orientation. While the authors can distinguish between a parallel vs. perpendicular orientation of this sequence relative to the plane of the membrane bilayer surface from their simulations, suggesting that previous studies indicated a transmembrane orientation for Drp1 is disingenuous and misleading. The term "transmembrane" should be removed or replaced, as it presents a wrong image.

      (3) Mutational analysis of W vs. F in membrane insertion of W-centered insertion motifs and vice versa: The PPM-based workflow suggests that F-centered sequences have the highest membrane insertion properties as opposed to W-centered ones. A W552F mutation in the WRGML sequence of Drp1 was, however, found to impair function. How do the authors rationalize this? A cross-mutational analysis of W vs. F in W-centered motifs and F-centered motifs is warranted.

    5. Author response:

      eLife Assessment

      This study provides a valuable advance in understanding how disordered proteins interact with cell membranes by identifying the sequence rules that enable aromatic residues to penetrate deeply into the membrane interior. The integration of complementary computational approaches, including molecular simulations, large-scale sequence analysis, and the development of an online prediction server, makes the work potentially impactful for the membrane protein and intrinsically disordered protein communities. The evidence supporting the main conclusions is generally convincing, although its transferability across diverse membrane compositions and its validity as a prediction tool for real protein-membrane systems remain to be further established.

      We thank the editors for recognizing our study as a valuable advance. This work lays a solid foundation for future developments to account for diverse membrane compositions and further refinements after additional experimental tests.

      Public review:

      Reviewer #1:

      A primary limitation is the heavy reliance on computational modeling. Training for AroMIP is generated using PPM rather than direct experimental measurements, and so the model may primarily reproduce PPM behavior rather than true membrane insertion thermodynamics. Moreover, all simulations use a single lipid composition (POPC:POPS:PIP<sub>2</sub> 70:25:5), but biological membranes vary substantially in cholesterol, cardiolipin, and acidic lipid content. Whether AroMIP's predictions transfer to diverse lipid environments remains untested. The 5% PIP<sub>2</sub> concentration used in the simulations is higher than that of a normal mammalian cell and may therefore overemphasize electrostatic contributions. Applicability beyond short 9-residue motifs is unclear, as longer-range interactions or secondary structure in full-length IDRs could modulate insertion in ways the current model does not capture. This could be considered for future development.

      The reviewer’s point on our reliance on PPM for training, a single lipid composition, and potential effects beyond a 9-residue motif is well taken. Regarding PPM, we chose it as the optimal compromise for high-throughput data. However, we complemented the high-throughput PPM data with experimental data on an initial set of 10 peptides. Moreover, we validate AroMIP on an additional 12 IDRs (intrinsically disordered regions; Table S2). On membrane composition, we now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). On potential effects beyond a 9-residue motif, we now add justification and note neglected factors for future developments (paragraph running from p. 19-20), as suggested by the reviewer.

      Reviewer #2:

      (1) Aromatic residues have been shown to partition preferentially to the headgroup region of the lipid bilayer. Most of the papers on this problem were published in the mid 1990s to early 2000s. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev. Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602-608; Killian & von Heijne, TIBS 2000, 25, 429-434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, none of these articles is cited.

      We have now citations to the Landolt-Marticorena paper and the von Heijne reviews (refs 25-27). The Doyle paper is not particularly relevant. As for the Fleming paper, we cited a 2016 JACS paper (original ref 27; now ref 30) that specifically dealt with aromatic residues.

      (2) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth".”

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is incorrect even in a qualitative sense.

      We have now revised these figures.

      Reviewer 3:

      (1) Membrane composition and lipid shape characteristics: The authors chose to use a model membrane bilayer of a distinct lipid composition, POPC: POPS: PI4,5P2 (70:25:5 molar ratio), for their all-atom simulations of the various model peptides. While this may be pertinent for some of these peptides, it is not for many, such as sequence 2 derived from Drp1, which preferentially binds target conical lipids such as cardiolipin (CL) and phosphatidic acid (PA). The rationale behind using PI4,5P2, which can induce positive membrane curvature when sequestered, versus CL and PA, which both induce negative membrane curvature, is not explained.

      We now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). In this Discussion paragraph, we also speculate that conical lipids, by promoting membrane defects, may facilitate membrane insertion.

      (2) Parallel vs. perpendicular peptide orientation of sequence 2 in peripheral Drp1-lipid interactions: On page 11, the authors state that their simulation results of sequence 2 derived from Drp1 "contrasts with a transmembrane orientation proposed by Mahajan et al." However, upon review, a transmembrane orientation for this region has never been proposed anywhere. Drp1 is a peripheral membrane protein that reversibly binds CL- and PA-containing membranes via its intrinsically disordered variable domain containing an aromatic-centered WRG motif. Indeed, the model presented in Figure 9 of Mahajan et al. displays a peripheral and parallel orientation of the transiently helical WRG-containing motif rather than a transmembrane (i.e., across the bilayer) orientation. While the authors can distinguish between a parallel vs. perpendicular orientation of this sequence relative to the plane of the membrane bilayer surface from their simulations, suggesting that previous studies indicated a transmembrane orientation for Drp1 is disingenuous and misleading. The term "transmembrane" should be removed or replaced, as it presents a wrong image.

      We have now deleted the sentence mentioning “transmembrane orientation”.

      (3) Mutational analysis of W vs. F in membrane insertion of W-centered insertion motifs and vice versa: The PPM-based workflow suggests that F-centered sequences have the highest membrane insertion properties as opposed to W-centered ones. A W552F mutation in the WRGML sequence of Drp1 was, however, found to impair function. How do the authors rationalize this? A cross-mutational analysis of W vs. F in W-centered motifs and F-centered motifs is warranted.

      AroMIP predicts a membrane insertion propensity of 0.782 for the WRGML sequence and a moderately higher propensity, 0.837, with a W552F mutation. This increase contradicts the experimental observation of a 3.6-fold increase in membrane binding affinity by Mahajan et al. We now speculate that the specific lipid, cardiolipin, as the reason for the discrepancy (p. 19, 3rd paragraph). This discrepancy provides a concrete example for the need to account for membrane composition in future developments.

    1. eLife Assessment

      This important study combines chromatin accessibility and genomic DNA sequence conservation data from low-coverage genome sequencing of related species (without assembly), for the in silico identification of cis-regulatory elements in large genomes. The approach and results are compelling and well supported by the experimental validations. The work will be of interest to researchers working in the field of gene regulation and evolution, particularly because the methodology proposed can be applied to a large variety of experimental organisms.

    2. Reviewer #1 (Public review):

      Summary:

      Forbes et al. developed an integrated approach to identify cis-regulatory elements (CREs) in the large (3.6 Gbp) genome of the crustacean Parhyale hawaiensis, addressing the challenge of pinpointing these regions among large regions of non-coding sequences. They combined ATAC-seq chromatin accessibility profiling (both bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis). Without assembling congener genomes, they mapped reads with low stringency to the P. hawaiensis reference, identifying about 55k conserved islands that overlap ATAC peaks more than expected by chance. This dual filter was used to select CRE candidates for transgenic reporter validation, yielding 6 functional elements (out of 11 tested) driving ubiquitous, neuronal, or muscle-specific expression, a major advance for non-model systems with large genomes.

      Strengths:

      Forbes et al. generated high-quality ATAC data across multiple scales. Using bulk ATAC-seq (from whole embryos, developing and adult legs), they identified tens of thousands of open chromatin peaks across the assembled P. hawaiensis large genome. Moreover, using single-nucleus ATAC-seq from adult legs, they could resolve differentially accessible chromatin profiles across over 15 cell types previously identified by scRNA-seq, enabling cell-type-specific candidate selection.

      Furthermore, their innovative low-coverage comparative genomics method mapped 0.46-6.4% of congener reads to P. hawaiensis without genome assembly, revealing hundreds of thousands of conserved non-coding islands, including about 55k showing conservation in all four species, far exceeding random expectation.

      Using the developed approach, the authors could validate 6 (out of 11 candidates) reporter constructs, driving robust ubiquitous and tissue-specific expression, succeeding where prior promoter-only screening failed and providing immediately useful genetic tools for the Parhyale community.

      Weaknesses:

      The primary limitation is that functional CRE testing was performed only in P. hawaiensis. While conservation maps are valuable resources, the manuscript lacks functional validation in congener species, limiting claims about broad applicability across related genomes/species.

      The approach also failed to validate developmental CREs. None of the candidates from combined ATAC and conservation filtering drove reporter expression matching endogenous patterns. The authors appropriately hypothesize technical limits (low expression) or biological factors (long-range enhancers, shadow enhancers).

      Overall Assessment:

      Forbes et al. fully succeed with their integrated approach to (1) generate an ATAC-seq atlas plus functional CRE discovery and (2) innovative low-coverage sequencing for conservation mapping in the large 3.6 Gbp genome of Parhyale hawaiensis. Their combination of ATAC-seq chromatin accessibility profiling (bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis), without congener genome assembly, drastically shrank the CRE search space. Using this approach, the authors could validate six out of 11 candidate transgenic reporters (ubiquitous, neuronal, and muscle-specific), where prior promoter-only screening failed.

      The low-coverage mapping innovation cuts cost and labour while snATAC-seq provides cell-type resolution, making these resources valuable for building new genetic and imaging tools in Parhyale.

      This compelling method also has the potential to enable labs with limited resources to identify and characterize regulatory elements in more non-model organisms, advancing our understanding of their evolution while establishing a scalable pipeline for large-genome systems.

    3. Reviewer #2 (Public review):

      The manuscript by Forbes, Skafida, Karapidaki et al. concerns the in silico identification of cis-regulatory elements (CREs) in large genomes using chromatin accessibility (ATAC-seq) and sequence conservation (genomic DNA sequencing) data. They exemplify this method by applying it to identify novel CREs in Parhyale hawaiensis, which they validated using reporter constructs.

      The results are convincing and are well supported by the data and validations. Identified CREs are valuable for researchers interested in the regulation of the expression of genes they control.

      The methodology on the whole is also valid, as suggested by the results and previous publications on various taxa. Sequence conservation, as stated by the authors, was long used as a method to identify regions of non-coding DNA with functional and evolutionary constraints. The same applies to ATAC-seq data, which has also been used as a proxy for functional regions in different animals such as sea urchins and amphioxus. The methodology proposed is likely to be successfully used by researchers working on a variety of experimental organisms.

      The authors do not use existing genome assemblies and use short-read sequencing to identify conserved regions, and while it is not conceptually novel, such an approach is becoming more and more viable and useful considering the recent advances in next-generation sequencing technology and the decrease in price of short-read sequencing.

      Two major weaknesses are:

      (1) The novelty of the approach and its advantages should be more explicitly stated.

      (2) The authors do not discuss in depth the strength of using a combination of two methods rather than either of the two, especially considering that previously known CREs do not overlap with conserved sequences.

    4. Reviewer #3 (Public review):

      Summary:

      Forbes et al. present a new approach for identifying cis-regulatory elements in large genomes. Using Parhyale hawaiensis, a crustacean with a large genome (~3.6 Gb, comparable in size to the human genome), the authors show that current methods for identifying cis-regulatory elements, effective in smaller genomes, are markedly inefficient in organisms with large genomes. To address this limitation, they combine bulk ATAC-seq and single-cell (sc) ATAC-seq to identify chromatin regions that are either ubiquitously accessible or specifically accessible in particular cell types. They further integrate comparative genomics across multiple Parhyale species (P. hawaiensis, P. aquilina, and P. darvishi), selected at appropriate phylogenetic distances (20-95 million years divergence), to pinpoint conserved open chromatin regions likely under functional constraint.

      Using this strategy, the authors predict a set of ubiquitous and cell-type-specific cis-regulatory elements. Importantly, they validate these predictions using rigorous transgenic reporter assays, convincingly demonstrating that their approach can successfully identify functional regulatory elements where previous methods had failed.

      Strengths:

      The approach introduced by Forbes et al. is conceptually straightforward, efficient, and readily transferable to other organisms. The validation experiments show not only that a substantial proportion of the predicted elements are functional, but also that the method is capable of identifying both ubiquitous and cell-type-specific regulatory elements. Given that the identification of regulatory regions remains a major bottleneck in understanding the molecular mechanisms underlying processes of development and regeneration, this work has the potential to make a significant impact in developmental and regeneration biology, particularly for studies involving non-model organisms with large genomes.

      An additional strength is the demonstration that only the genome of the focal species requires high-quality sequencing and assembly. In contrast, species used solely for comparative analysis can be sequenced at low coverage without assembly, substantially reducing costs and increasing the accessibility of the approach.

      Weaknesses:

      While the method is effective in identifying regulatory elements that are active ubiquitously or in differentiated cell types, it failed in detecting elements associated with developmentally regulated genes. This may be due to trivial reasons, such as a very low level of expression of the selected genes. However, as acknowledged by the authors, it may also indicate inherent challenges in identifying regulatory elements associated with developmentally dynamic gene regulation, compared to those associated with genes expressed in differentiated cell types.

      A second limitation, also acknowledged by the authors, is the absence of chromatin conformation capture data, which would help link distal regulatory elements to their target genes. This limitation may be particularly relevant for developmentally regulated genes, where long-range regulatory interactions may be critical.

      Addressing these limitations will be an important direction for future work. Nonetheless, the approach as presented in this manuscript represents a key contribution that sets the stage for further methodological advances in the identification of cis-regulatory elements in large genomes.

    1. eLife Assessment

      This is an important study showing the interaction of the endoplasmic reticulum (ER)-resident tyrosine phosphatase PTP1B with the developing phagocytic cup in macrophages, and its role in inhibiting microbicidal superoxide production. The authors show convincing evidence that PTP1B interacts with Syk, a plasma membrane tyrosine kinase that plays an essential role in phagocytosis, and that ablation of PTP1B increases superoxide production and Syk phosphorylation without affecting phagocytosis. Further evidence suggests that PTP1B may inhibit a Syk/Shc1/NOX2 axis; however, robust demonstration of the proposed chain of events and of the actual role of ER-plasma membrane contact sites in the PTP1B-dependent downregulation of NOX2 activity will require additional experimental evidence. The integration of advanced imaging methods to study contact site formation with functional assays related to phagocytosis and signaling is inspiring.

    2. Joint Public Review:

      Summary:

      This study uses state-of-the-art imaging approaches to show that membrane contact site (MCS) markers and the ER-resident tyrosine phosphatase PTP1B accumulate on phagocytic membranes within actin-devoid zones during frustrated phagocytosis in RAW264.7 macrophages. The authors convincingly show that PTP1B interacts with Syk, an Fcγ receptor-associated tyrosine kinase that plays a critical role in phagocytosis, and that ablation of PTP1B results in hyperphosphorylation of Syk and increased superoxide production, without impacting phagocytic efficiency. Using a phosphoproteomic approach, the authors identify the adaptor protein Shc1 as a strongly phosphorylated protein during stimulation of immunoglobulin receptors by aggregated IgG. In the absence of PTP1B, the authors demonstrate an increased interaction between Shc1 and the NADPH oxidase NOX2 subunit p47phox, suggesting that PTP1B controls superoxide production by inhibiting a Syk-Shc1-NOX2 axis.

      Strengths:

      This is a well-reasoned and cogently developed study that uses contemporary methods, including high-quality TIRF microscopy combined with MAPPER (Membrane-Attached Peripheral ER) or SPLICS (split-GFP-based contact site sensors), to describe how membrane contact site markers and the ER-resident tyrosine phosphatase PTP1B accumulate in the phagocytic cup as cortical actin depolymerizes. The genetic data also convincingly show that PTP1B ablation increases Syk and Shc1 phosphorylation, enhances the Shc1/p47phox interaction, and elevates superoxide production, whereas depletion of Shc1 reduces superoxide levels. Overall, the work outlines an interesting interplay between membrane contact sites, signaling, and the phagocytic machinery of broad interest.

      Weaknesses:

      While the authors indicate that the PTP1B phosphatase downregulates superoxide production via the Syk-Shc1-NOX2 axis and present a summary model depicting the proposed sequence of events, the supporting data are currently mostly circumstantial. For example, although it is clear that PTP1B depletion increases superoxide production as well as Syk and Shc1 phosphorylation in vivo, there are no data directly demonstrating that the effects of PTP1B depletion on superoxide production require enhanced Syk or Shc1 phosphorylation. Likewise, although PTP1B depletion increases the interaction between Shc1 and p47phox, a soluble component of NOX2, there is no compelling demonstration that superoxide production in PTP1B-depleted cells truly depends on the NOX2 complex or on the Shc1/p47phox interaction.<br /> In addition, while the authors elegantly demonstrate the formation of ER-PM contact sites during frustrated phagocytosis within the actin clearance zone, as well as the localization of the PTP1B phosphatase in the same region, it remains unclear whether the presence of the phosphatase at membrane contact sites is required for its regulatory effect on superoxide production.

      Finally, it would be interesting to investigate these phenomena in other macrophage cell lines and perhaps also in more physiological contexts than frustrated phagocytosis. This would help evaluate the broader generalizability of the results and conclusions.

    1. eLife Assessment

      This important study combined careful computational modeling, a large patient sample, and replication in an independent general population sample to provide convincing evidence in support of a computational account of a difference in risk-taking between people who have attempted suicide and those who have not. It is proposed that this difference reflects a general change in the approach to risky (high-reward) options and a lower emotional response to certain rewards. While the findings advance our understanding of cognitive mechanisms at the group level, the observation that computational phenotype is predictive of suicidal behavior only in the clinical sample and not in the online sample limits its applicability for individual prediction, early detection and prevention of suicidality.

    2. Reviewer #1 (Public review):

      Summary:

      The authors use a gambling task with momentary mood ratings from Rutledge et al. and compare computational models of choice and mood to identify markers of decisional and affective impairments underlying risk-prone behavior in adolescents with suicidal thoughts and behaviors (STB). The results show that adolescents with STB show enhanced gambling behavior (choosing the gamble rather than the sure amount), and this is driven by a bias towards the largest possible win rather than insensitivity to possible losses. Moreover, this group shows a diminished effect of receiving a certain reward (in the non-gambling trials) on mood. The results were replicated in a general online sample where participants were divided into groups with or without STB based on their self-report of suicidal ideation on one question in the Beck Depression Inventory self-report instrument. The authors suggest, therefore, that adolescents diagnosed with depression or anxiety with decreased sensitivity to certain rewards may need to be monitored more closely for STB due to their increased propensity to take risky decisions aimed at (expected) gains (such as relief from an unbearable situation through suicide) regardless of the potential losses. However, such a result was only found in the clinical sample and cannot be generalized more broadly based on the current findings.

      Strengths:

      ● The study uses a previously validated task design and replicates previously found results through well-explained model-free and model-based analyses.

      ● Sampling of adolescents at high risk can help target early preventative diagnoses and treatments for suicide.

      ● Replication of the results in an online cohort increases confidence in the findings.

      ● The models considered for comparison are thorough and well-motivated. The chosen models allow for teasing apart which decision and mood sensitivity parameters relate to risky decision-making across groups based on their hypotheses.

      ● Novel finding of mood (in)sensitivity to non-risky rewards and its relationship with risk behavior in STB.

      Weaknesses:

      ● Sample size of 25 for S- group is low-powered, which is explicitly mentioned as a study limitation.

      ● Modeling in the mediation analysis focused on predicting risk behavior in this task from the model-derived bias for gains and suicidal symptom scores. Thus, the implications of this work are more relevant to a basic-science understanding of the etiology of suicidal behavior than they are useful as a predictor of suicidal behavior, and it is not clear that a psychiatrist or psychologist could use this task to potentially determine who is at higher risk of attempting suicide and must be more closely monitored. Indeed, relationships between task parameters and behavior and suicidal behavior was limited to the clinical sample with a diagnosis of depression or anxiety disorder, and did not extend to the online sample. Therefore, the claim that these findings provide "computational markers for general suicidal tendency among adolescents" is unwarranted.

    3. Reviewer #2 (Public review):

      Summary:

      This article addresses a very pertinent question - what are the computational mechanisms underlying risky behaviour in patients having attempted suicide. In particular, it is impressive how the authors find a broad behavioral effect whose mechanisms they can then explain and refine through computational modeling. This work is important because currently, beyond previous suicide attempts, there has been a lack of predictive measures. This study is the first step towards that: understanding the cognition on a group level. Before then being able to include it in future predictive studies (based on the cross-sectional data, this study by itself cannot assess the predictive validity of the measure).

      Strengths:

      - Large sample size

      - Replication of their own findings

      - Well-controlled task with measures of behaviour and mood + precise and well-validated computational modeling

      Questions, based on revised manuscript and replies to other reviewers:

      (1) Replies to reviewers in general: Bayes Factors have been added, it would be good to also use common verbal terms to describe them (e.g. 'anecdotal', 'moderate' etc). For example, my reading of table S8 would be that for gambling rate there is only anecdotal evidence that it does not relate to PSWQ, BDI, and moderate evidence it does not relate to TAI.

      (2) Reply to reviewer 1 Q2 (Predicting STB):

      For the regression predicting suicidal ideation, it seems to me that what you did was a regression STB ~ gambling behaviour + approach + mood? Could you clarify? I had expected as a test of whether the task can predict STB risk something slightly different - a cross-validation (LOO or maybe 5-fold in the large sample): STB ~ gambling behaviour + approach [parameter from model] + mood [parameter from model]; and then computing in the left out participants: predicted STB. Then checking correlation between STB and predicted STB. This would allow testing whether the diverse task measures together predict STB (with the caveat, that it's cross-validated, rather than hold-out sample, unless you could train on one sample (in lab) and test on the other (online).

      (3) Reply to reviewer 2 Q1 (parameter recovery): I'm looking at S3, it seems to still show only the scatter plots and not the correlation matrices, which are now added as text notes. Can you actually show these matrices? An off-diagonal correlation of 0.63 appears quite high. I think it needs to be discussed exactly which parameters those are, and whether that impacts the interpretation of the results.

      (4) Reply to reviewer 3 Q3 (mood model): I would have imagined that the response would involve changing the mood equations (equation 8 main text) to include a term for whether the participant gambled or not, independent of the gamble value.

    4. Reviewer #3 (Public review):

      This manuscript investigates computational mechanisms underlying increased risk-taking behavior in adolescent patients with suicidal thoughts and behaviors. Using a well-established gambling task that incorporates momentary mood ratings and previously established computational modeling approaches, the authors identify particular aspects of choice behavior (which they term approach bias) and mood responsivity (to certain rewards) that differ as a function of suicidality. The authors replicate their findings on both clinical and large-scale non-clinical samples.

      The main problem, however, is that the results do not seem to support a specific conclusion with regard to suicidality. The S+ and S- groups differ substantially in the severity of symptoms, as can be seen by all symptom questionnaires and the baseline and mean mood, where S- is closer to HC than it is to S+. The main analyses control for illness duration and medication but not for symptom severity. The supplementary analysis in Figure S11 is insufficient as it mistakes the absence of evidence (i.e., p > 0.05) for evidence of absence. Therefore, the results do not adequately deconfound suicidality from general symptom severity.

      The second main issue is that the relationship between an increased approach bias and decreased mood response to CR is conceptually unclear. In this respect, it would be natural to test whether mood responses influence subsequent gambling choices. This could be done either within the model by having mood moderate the approach bias or outside the model using model-agnostic analyses.

      Additionally, there is a conceptual inconsistency between the choice and mood findings that partly results from the analytic strategy. The approach bias is implemented in choice as a categorical value-independent effect, whereas the mood responses always scale linearly with the magnitude of outcomes. One way to make the models more conceptually related would be to include a categorical value-independent mood response to choosing to gamble/not to gamble.

      The manuscript requires editing to improve clarity and precision. The use of terms such as "mood" and "approach motivation" is often inaccurate or not sufficiently specific. There are also many grammatical errors throughout the text.

      Claims of clinical relevance should be toned down, given that the findings are based on noisy parameter estimates whose clinical utility for the treatment of an individual patient is doubtful at best.

      Comments on revisions:'

      The authors adequately addressed my comments and I find the manuscript substantially strengthened.

    5. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study combined careful computational modeling, a large patient sample, and replication in an independent general population sample to provide a computational account of a difference in risk-taking between people who have attempted suicide and those who have not. It is proposed that this difference reflects a general change in the approach to risky (high-reward) options and a lower emotional response to certain rewards. Evidence for the specificity of the effect to suicide, however, is incomplete, which would require additional analyses.

      We thank the editors and reviewers for this important assessment. Based on clinical interviews, we included patients with and without suicidality (S<sup>+</sup> and S<sup>-</sup> groups). However, in line with suicidal-related literature (e.g., Tsypes et al., 2024), two groups also differed substantially in the severity of symptoms (see Table 1). To address the request for evidence on specificity to suicidality beyond general symptom severity, we performed separate linear regressions to explain in gambling behaviour, value-insensitive approach parameter (β<sub>gain</sub>), and mood sensitivity to certain rewards (β<sub>CR</sub>) with group as a predictor (1 for S<sup>+</sup> group and 0 for S<sup>-</sup> group) and scores for anxiety and depression as covariates. Results remained significant after controlling anxiety and depression (ps < 0.027; Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) on the clinical questionnaire to extract the orthogonal components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. We then performed linear regressions using these components as covariates to control for anxiety and depression. Our main results remained significant (ps < 0.027; Table S9). We believe that these analyses provide evidence that the main effects on gambling and on mood were specific to suicide.

      Moreover, as Reviewer 3 pointed out, these “absence of evidence” cannot provide insights of “evidence of absence”. Although we median-split patients by the scores of general symptoms (e.g., depression and anxiety-related questionnaires) and verified no significant differences in these severities (Figure S11), we additionally conducted Bayesian statistics in gambling behavior, value-insensitive approach parameter, and mood sensitivity to certain rewards. BF<sub>01</sub> is a Bayes factor comparing the null model (M<sub>0</sub>) to the alternative model (M<sub>1</sub>), where M<sub>0</sub> assumes no group difference. BF<sub>01</sub> > 1 indicates that evidence favors M<sub>0</sub>. As can be seen in Table S7, most results supported null hypothesis, suggesting that general symptoms of anxiety and depression overall did not influence our main results. Overall, we believe that these analyses provide compelling evidence for the specificity of the effect to suicide, above and beyond depression and anxiety.

      Beyond these specific findings, this work highlights the broader utility of computational modelling and mood to better understand behavioral effect, showing how to use both mood and choice data to better comprehend a psychiatric issue.

      Please see Tables S7, S8, S9 and our revisions below:.

      Page 17:

      “Within patients, this group effect on gambling rate remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.024; also see Figure S11, Table S7 and Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, (ps < 0.001), we performed Principal Components Analysis (PCA) to extract main components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. To further control for anxiety and depression, linear regression using these components as covariates revealed that the group effect on gambling rate remained significant (p = 0.024; Table S9).”

      Pages 18-19:

      “Within patients, this group effect on the approach parameter remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.027; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on approach parameter remained significant (p = 0.027; Table S9).”

      Page 21:

      “Within patients, this group effect on βCR remained significant after controlling for gambling rate, earnings, mood-related outcome effect, mood drift effect, sex, illness duration, family history, diagnosis, and various medications use (ps < 0.032), as well as general symptoms (e.g., depression and anxiety; p = 0.001; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on this mood parameter remained significant (p = 0.001; Table S9).”

      Page 27:

      “Beyond these specific findings, this work highlights the broader utility of computational modelling and mood to better understand behavioral effect, showing how to use both mood and choice data to better comprehend a psychiatric issue.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors use a gambling task with momentary mood ratings from Rutledge et al. and compare computational models of choice and mood to identify markers of decisional and affective impairments underlying risk-prone behavior in adolescents with suicidal thoughts and behaviors (STB). The results show that adolescents with STB show enhanced gambling behavior (choosing the gamble rather than the sure amount), and this is driven by a bias towards the largest possible win rather than insensitivity to possible losses. Moreover, this group shows a diminished effect of receiving a certain reward (in the non-gambling trials) on mood. The results were replicated in an undifferentiated online sample where participants were divided into groups with or without STB based on their self-report of suicidal ideation on one question in the Beck Depression Inventory self-report instrument. The authors suggest, therefore, that adolescents with decreased sensitivity to certain rewards may need to be monitored more closely for STB due to their increased propensity to take risky decisions aimed at (expected) gains (such as relief from an unbearable situation through suicide), regardless of the potential losses.

      Strengths:

      (1) The study uses a previously validated task design and replicates previously found results through well-explained model-free and model-based analyses.

      (2) Sampling choice is optimal, with adolescents at high risk; an ideal cohort to target early preventative diagnoses and treatments for suicide.

      (3) Replication of the results in an online cohort increases confidence in the findings.

      (4) The models considered for comparison are thorough and well-motivated. The chosen models allow for teasing apart which decision and mood sensitivity parameters relate to risky decision-making across groups based on their hypotheses.

      (5) Novel finding of mood (in)sensitivity to non-risky rewards and its relationship with risk behavior in STB.

      Weaknesses:

      (1) The sample size of 25 for the S- group was justified based on previous studies (lines 181-183); however, all three papers cited mention that their sample was low powered as a study limitation.

      We thank the Reviewer for rising this concern. We agree that the sample size for S<sup>-</sup> group (n=25) is modest, and the prior studies we cited also acknowledged limited power. We wanted to point out that we obtained a comparable sample size to a prior study. In the revision, we therefore updated the section to justify this sample size in which we acknowledge the limited power of our study in the limitation section. Please see our clarification below:

      Page 32:

      “Third, despite replicating our main results in an independent dataset (n=747), the modest S<sup>-</sup> subgroup size (n=25) has a limited statistical power.”

      (2) Modeling in the mediation analysis focused on predicting risk behavior in this task from the model-derived bias for gains and suicidal symptom scores. However, the prediction of clinical interest is of suicidal behaviors from task parameters/behavior - as a psychiatrist or psychologist, I would want to use this task to potentially determine who is at higher risk of attempting suicide and therefore needs to be more closely watched rather than the other way around (predicting behavior in the task from their symptom profile). Unfortunately, the analyses presented do not show that this prediction can be made using the current task. I was left wondering: is there a correlation between beta_gain and STB? It is also important to test for the same relationships between task parameters and behavior in the healthy control group, or to clarify that the recommendations for potential clinical relevance of these findings apply exclusively to people with a diagnosis of depression or anxiety disorder. Indeed, in line 672, the authors claim their results provide "computational markers for general suicidal tendency among adolescents", but this was not shown here, as there were no models predicting STB within patient groups or across patients and healthy controls.

      Thank you for these thoughtful comments. Our study focuses on why adolescent patients with suicidality have increased risk behavior, aiming to provide a mechanism-based target for suicide prevention. Therefore, our dependent variable in the mediation model was gambling behavior. We also agree that the clinically relevant question is whether suicidality can be predicted from task-derived behavior/parameters. We thus used risky behavior and the potential mental parameters to predict STB. Linear regressions showed that gambling behavior, as well as the value-insensitive approach parameter, can predict suicidal symptom scores among patients (former: β = 9.189, t = 2.004, p = 0.048; latter: β = 5.587, t = 2.890, p = 0.005). In healthy controls, these predictions failed (gambling behavior: β = 1.471, t = 0.825, p = 0.411; approach: β = 0.874, t = 1.178, p = 0.241). These results suggest that clinical relevance of these findings apply exclusively to people with a diagnosis of depression or anxiety disorder. We found same patterns for the mood parameter (mood sensitivity to certain rewards: patients: β = -28.706, t = -2.801, p = 0.006; healthy controls: β = -2.204, t = -0.528, p = 0.599). In sum, we believe that our statement of “computational markers for general suicidal tendency among adolescents” is reasonable now. Please see our revisions below:

      Page 17:

      “Furthermore, linear regression showed that gambling rate can predict the current suicidal ideation score (BSI-C, β = 9.189, t = 2.004, p = 0.048) among patients, but not among HC (β = 1.471, t = 0.825, p = 0.411), suggesting that gambling behavior has patient-specific predictive utility for suicidal symptoms.”

      Page 19:

      “Furthermore, linear regression showed that approach parameter can predict the current suicidal ideation score (β = 5.587, t = 2.890, p = 0.005) among patients, but not among HC (β = 0.874, t = 1.178, p = 0.241), suggesting that value-insensitive approach parameter has patient-specific predictive utility for suicidal symptoms.”

      Page 21:

      “Furthermore, linear regression showed that mood sensitivity to CR can predict the current suicidal ideation score (β = -28.706, t = -2.801, p = 0.006) among patients, but not among HC (β = -2.204, t = 0.528, p = 0.599), suggesting that mood sensitivity to CR has patient-specific predictive utility for suicidal symptoms.”

      (3) The FDR correction for multiple comparisons mentioned briefly in lines 536-538 was not clear. Which analyses were included in the FDR correction? In particular, did the correlations between gambling rate and BSI-C/BSI-W survive such correction? Were there other correlations tested here (e.g., with the TAI score or ERQ-R and ERQ-S) that should be corrected for? Did the mediation model survive FDR correction? Was there a correction for other mediation models (e.g., with BSI-W as a predictor), or was this specific model hypothesized and pre-registered, and therefore no other models were considered? Did the differences in beta_gain across groups survive FDR when including comparisons of all other parameters across groups? Because the results were replicated in the online dataset, it is ok if they did not survive FDR in the patient dataset, but it is important to be clear about this in presenting the findings in the patient dataset.

      Thank you for raising the important issue of multiple testing and for asking us to clarify exactly which tests were covered by the FDR procedure. In the clinical dataset we conducted a large number of inferential tests (χ<sup>2</sup>, t-tests, ANOVAs, regressions) spanning: (i) group differences in demographic/clinical characteristics; (ii) sanity checks (e.g., anxiety/depression questionnaires); (iii) primary hypotheses (e.g., group differences in risky behavior); (iv) model-based analyses (parameter checks and between-group contrasts); and (v) control/sensitivity analyses. Post-hoc t-tests were performed only when the three-group ANOVA was significant. This yielded >150 p-values. FDR was applied using all these p-values. Please see Supplementary Note 8.

      (4) There is a lack of explicit mention when replication analyses differ from the analyses in the patient sample. For instance, the mediation model is different in the two samples: in the patient sample, it is only tested in S+ and S- groups, but not in healthy controls, and the model relates a dimensional measure of suicidal symptoms to gambling in the task, whereas in the online sample, the model includes all participants (including those who are presumably equivalent to healthy controls) and the predictor is a binary measure of S+ versus S- rather than the response to item 9 in the BDI. Indeed, some results did not replicate at all and this needs to be emphasized more as the lack of replication can be interpreted not only as "the link between mood sensitivity to CR and gambling behavior may be specifically observable in suicidal patients" (lines 582-585) - it may also be that this link is not truly there, and without a replication it needs to be interpreted with caution.

      Thank you for these important comments. This study focused on cognitive and affective computational mechanisms underlying increased risky behavior in STB. Accordingly, we compared patients with STB (S<sup>+</sup>) with patients without STB (S<sup>-</sup>) and healthy controls (HC) to examine the effects of STB on risky behavior. Therefore, group comparison, instead of dimensional measure of suicidal symptoms by Beck Scale for Suicidal Ideation, can answer our research questions directly.

      To enhance consistency between the clinical and replication datasets, we included all participants in each dataset when performing the mediation analysis. Given that S<sup>-</sup> and HC did not differ in gambling behavior or the approach parameter in the clinical dataset, we merged these two groups. In the replication dataset, to mirror the S<sup>+</sup> vs. S<sup>-</sup> contrast used clinically, we categorized the general sample into S<sup>+</sup> and S<sup>-</sup> based on BDI item 9. The mediation results remained significant in both datasets (the clinical dataset: a×b = 0.321, 95% CI = [0.070, 0.549], p = 0.016; the replication dataset: a × b = 0.143, 95% CI = [0.016, 0.288], p = 0.031), suggesting that STB is associated with increased risk behavior via stronger approach motivation.

      We also acknowledge the non-replication of the correlation between gambling behavior and mood sensitivity to certain rewards in the online sample. While this pattern might indicate that the link is specific to suicidal patients, it may also reflect sample-specific or unstable effects; thus, we now state this explicitly and interpret the finding with caution. Please see our revisions below:

      Page 15:

      “We next verified our results in an independent dataset, including the same task and BDI questionnaire in 747 general participants (500 females; age: 20.90±2.41)[46]. One item in BDI involves the measurement of STB. In item 9 of BDI, participants chose one option that describes them best: Option 1, “I don't have any thoughts of killing myself.”; Option 2, “I have thoughts of killing myself, but I would not carry them out.”; Option 3, “I would like to kill myself.”; Option 4, “I would kill myself if I had the chance.”. In line with the current definition of S<sup>+</sup>/S<sup>-</sup> in the clinical dataset, we identified S<sup>+</sup> group as choosing Option 2, 3, or 4, while participants selecting Option 1 were categorized as S<sup>-</sup> group.”

      Page 19:

      “Given significant correlations between group, approach parameter, and gambling rate for gain trials (ps < 0.017), we further conducted a mediation analysis with the assumption of the mediating effect of approach motivation of suicidality on the risk behavior. Given that we aimed to test the effect of STB, with S<sup>-</sup> and HC as controls, and given that S<sup>-</sup> and HC did not differ in gambling behavior or in the approach parameter, we merged these two groups for the mediation analysis. Results supported our hypothesis (a×b = 0.321, 95% CI = [0.070, 0.549], p = 0.016; Figure 2C), confirming that suicidal thoughts and behavior increase risk behavior through stronger approach motivation.”

      Page 26:

      “However, we did not observe any significant correlation between mood sensitivity to CR and gambling behavior (ps > 0.389), which suggests that the link between mood sensitivity to CR and gambling behavior may be specifically observable in suicidal patients. Alternatively, this non-replicated result may also reflect sample-specific or unstable effects, which needs to be interpreted with caution.”

      (5) In interpreting their results, the authors use terms such as "motivation" (line 594) or "risk attitude" (line 606) that are not clear. In particular, how was risk attitude operationalized in this task? Is a bias for risky rewards not indicative of risk attitude? I ask because the claim is that "we did not observe a difference in risk attitude per se between STB and controls". However, it seems that participants with STB chose the risky option more often, so why is there no difference in risk attitude between the groups?

      Thank you for pointing out the ambiguity. In our manuscript, “motivation” and “risk attitude” are defined at the computational level. Following prior work with this task Rutledge et al., (2015, 2016), we decompose observed gambling into (i) value-dependent valuation parameters that capture risk attitude (e.g., risk aversion and loss aversion, which scale the subjective value of outcomes), and (ii) value-insensitive, valence-dependent biases that capture approach/avoidance motivation. Accordingly, a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups which is what we observe for S<sup>+</sup> vs. controls. We have clarified this point in the computational modeling section.

      Pages 12-13:

      “Please note that a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups. Risk attitude is indeed conceptualized in economics as the curvature of the utility function (i.e., the subjective value) of the objective outcomes, with concave curves associated with risk aversion, and convex curves associated with risk seeking [54,56]. By contrast, the approach or avoidance bias apply to all the value. A possible interpretation of the approach bias is that participant approach the option with the highest possible gain (the lottery) in the gain frame; the avoidance bias would then reflect a tendency to systematically avoid the highest potential losses (the lottery) in the loss frame.”

      Reviewer #2 (Public review):

      Summary:

      This article addresses a very pertinent question: what are the computational mechanisms underlying risky behaviour in patients who have attempted suicide? In particular, it is impressive how the authors find a broad behavioural effect whose mechanisms they can then explain and refine through computational modeling. This work is important because, currently, beyond previous suicide attempts, there has been a lack of predictive measures. This study is the first step towards that: understanding the cognition on a group level. This is before being able to include it in future predictive studies (based on the cross-sectional data, this study by itself cannot assess the predictive validity of the measure).

      Strengths:

      (1) Large sample size.

      (2) Replication of their own findings.

      (3) Well-controlled task with measures of behaviour and mood + precise and well-validated computational modeling.

      Weaknesses:

      I can't really see any major weakness, but I have a few questions:

      (1) I can see from the parameter recovery that the parameters are very well identified. Is it surprising that this is the case, given how many parameters there are for 90 trials? Could the authors show cross-correlations? I.e., make a correlation matrix with all real parameters and all fitted parameters to show that not only the diagonal (i.e., same data is the scatter plots in S3) are high, but that the off-diagonals are low.

      Thank you for raising these thoughtful concerns. The current task consisted of 90 choices and 36 mood ratings. There were 5 choice parameters and 4 mood parameters. The apparently strong identifiability is not unexpected, as 90 choice trials and 36 mood ratings are comparable to those in prior computational modeling literature (Blain & Rutledge, 2022).

      As suggested, we computed cross-scorrelations between all generating (“true”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery. Please see Supplementary Pages 2-3.

      “Parameter recovery: Figure S3 shows good parameter recovery for both choice and mood winning model (choice: rs > 0.91, ps < 0.001; intraclass coefficients > 0.78; mood: rs > 0.90, ps < 0.001; intraclass coefficients > 0.86). Moreover, we computed cross-correlations between all generating (“true”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery.”

      Page 10:

      “The numbers of choice trials and mood ratings were comparable to those in prior computational modeling studies [34,35].”

      (2) Could the authors clarify the result in Figure 2B of a correlation between gambling rate and suicidal ideation score, is that a different result than they had before with the group main effect? I.e., is your analysis like this: gambling rate ~ suicide ideation + group assignment? (or a partial correlation)? I'm asking because BSI-C is also different between the groups. [same comment for later analyses, e.g. on approach parameter].

      Thank you for pointing out the lack of clarity. We performed group difference analysis and correlation of suicidal ideation analysis, separately. We first performed group difference analysis to test our hypothesis of STB effects. We then conducted correlational analysis to further specify our findings.

      (3) The authors correlate the impact of certain rewards on mood with the % gambling variable. Could there not be a more direct analysis by including mood directly in the choice model?

      Thank you for this insightful suggestion. As suggested, we tried to integrate mood into choice models by adding mood bias component(s) in line with previous literature (Vinckier et al., 2018). The first model (mcM1) assumes that mood biases choice, building on cM3 (the winning choice model). cmM2 further separated the mood bias parameter into two components according to participants’ choices.

      However, model comparison using BIC supported cM3 (Table S6), that is, without consideration of mood in choice modeling. This can be due to the lack of block design in our experimental design unlike e.g., Vinckier et al., (2018) and Eldar & Niv, (2015). Please see Supplementary Note 6.

      (4) In the large online sample, you split all participants into S+ and S-. I would have imagined that instead, you would do analyses that control for other clinical traits. Or, for example, you have in the S- group only participants who also have high depression scores, but low suicide items.

      Thank you for this insightful suggestion. Following prior suicide-related literature (Tsypes et al., 2024), we controlled for depression by including them as covariates. Note that depression scores were derived from our established bifactor model (Wang et al., 2025), which decomposed depression from the anxiety. These results remained largely significant (ps ≤ 0.050), except a marginally significant effect of group on gambling behavior (p = 0.059). Despite a trend, this effect with covariates of depression-related questionnaires is strong in our clinical cohort (p = 0.024; Table S8). This suggests that the link between suicidality and risky behavior persists above and beyond general depressive symptoms.

      Please see our clarifications below:

      Page 26:

      “After controlling for depression severity using our established bifactor model (see ref 60 for details), these results remained significant (ps ≤ 0.050), except a marginally significant effect of group on gambling behavior (p = 0.059). Despite a trend, this effect with covariates of depression-related questionnaires is strong in our clinical cohort (p = 0.024; Table S8). This suggests that the link between suicidality and risky behavior persists above and beyond general depressive symptoms.”

      Reviewer #3 (Public review):

      This manuscript investigates computational mechanisms underlying increased risk-taking behavior in adolescent patients with suicidal thoughts and behaviors. Using a well-established gambling task that incorporates momentary mood ratings and previously established computational modeling approaches, the authors identify particular aspects of choice behavior (which they term approach bias) and mood responsivity (to certain rewards) that differ as a function of suicidality. The authors replicate their findings on both clinical and large-scale non-clinical samples.

      (1) The main problem, however, is that the results do not seem to support a specific conclusion with regard to suicidality. The S+ and S- groups differ substantially in the severity of symptoms, as can be seen by all symptom questionnaires and the baseline and mean mood, where S- is closer to HC than it is to S+. The main analyses control for illness duration and medication but not for symptom severity. The supplementary analysis in Figure S11 is insufficient as it mistakes the absence of evidence (i.e., p > 0.05) for evidence of absence. Therefore, the results do not adequately deconfound suicidality from general symptom severity.

      Thank you for this important comment. Based on clinical interviews, we included patients with and without suicidality (S<sup>+</sup> and S<sup>-</sup> groups). However, in line with suicidal-related literature (e.g., Tsypes et al., 2024), two groups also differed substantially in the severity of symptoms (see Table 1). To address the request for evidence on specificity to suicidality beyond general symptom severity, we performed separate linear regressions to explain in gambling behaviour, value-insensitive approach parameter (β<sub>gain</sub>), and mood sensitivity to certain rewards (β<sub>CR</sub>) with group as a predictor (1 for S<sup>+</sup> group and 0 for S<sup>-</sup> group) and scores for anxiety and depression as covariates. Results remained significant after controlling anxiety and depression (ps < 0.027; Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) on the clinical questionnaire to extract the orthogonal components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. We then performed linear regressions using these components as covariates to control for anxiety and depression. Our main results remained significant (ps < 0.027; Table S9). We believe that these analyses provide evidence that the main effects on gambling and on mood were specific to suicide.

      As pointed out, these “absence of evidence” cannot provide insights of “evidence of absence”. Although we median-split patients by the scores of general symptoms (e.g., depression and anxiety-related questionnaires) and verified no significant differences in these severities (Figure S11), we additionally conducted Bayesian statistics in gambling behavior, value-insensitive approach parameter, and mood sensitivity to certain rewards. BF<sub>01</sub> is a Bayes factor comparing the null model (M<sub>0</sub>) to the alternative model (M<sub>1</sub>), where M<sub>0</sub> assumes no group difference. BF<sub>01</sub> > 1 indicates that evidence favors M<sub>0</sub>. As can be seen in Table S7, most results supported null hypothesis, suggesting that general symptoms of anxiety and depression overall did not influence our main results. Overall, we believe that these analyses provide compelling evidence for the specificity of the effect to suicide, above and beyond depression and anxiety.

      Please see Table S7, S8 &S9 and our revisions below.

      Page 17:

      “Within patients, this group effect on gambling rate remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.024; also see Figure S11, Table S7 and Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) to extract main components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. To further control for anxiety and depression, linear regression using these components as covariates revealed that the group effect on gambling rate remained significant (p = 0.024; Table S9).”

      Pages 18-19:

      “Within patients, this group effect on the approach parameter remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.027; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on approach parameter remained significant (p = 0.027; Table S9).”

      Page 21:

      “Within patients, this group effect on βCR remained significant after controlling for gambling rate, earnings, mood-related outcome effect, mood drift effect, sex, illness duration, family history, diagnosis, and various medications use (ps < 0.032), as well as general symptoms (e.g., depression and anxiety; p = 0.001; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on this mood parameter remained significant (p = 0.001; Table S9).”

      (2) The second main issue is that the relationship between an increased approach bias and decreased mood response to CR is conceptually unclear. In this respect, it would be natural to test whether mood responses influence subsequent gambling choices. This could be done either within the model by having mood moderate the approach bias or outside the model using model-agnostic analyses.

      Thank you for this important suggestion. As suggested, one interesting question was whether mood responses influence subsequent gambling choices and how to model them. First, we median-split mood responses (except the final rating) to compare gambling rate. Results showed a trend for less gambling rate in higher mood (t = -1.971, p = 0.050). However, there was no significant group difference (F = 0.680, p = 0.507). Second, with the assumption that mood biases choice, we constructed mcM1 based on cM3 (the winning choice model). Based on our finding of the negative correlation between mood sensitivity to certain rewards and gambling rate in S<sup>+</sup>, we separated β<sub>Mood</sub> parameter into β<sub>Mood-CR</sub> and β<sub>Mood-GR</sub> (cmM2). Model comparison using BIC supported cM3 (Table S6), that is, without consideration of mood in choice modeling. This can be due to the lack of block design in our experimental design unlike e.g., Vinckier et al., (2018) and Eldar & Niv, (2015). Please see Supplementary Note 6.

      (3) Additionally, there is a conceptual inconsistency between the choice and mood findings that partly results from the analytic strategy. The approach bias is implemented in choice as a categorical value-independent effect, whereas the mood responses always scale linearly with the magnitude of outcomes. One way to make the models more conceptually related would be to include a categorical value-independent mood response to choosing to gamble/not to gamble.

      We apology for the unclear statement. The approach bias is implemented in choice as a continuous value-independent effect, ranging from -1 to 1.

      It was true that the mood responses always scale with the magnitude of outcomes, since mood ratings were request after the outcomes. Therefore, mood parameters and the approach bias were both continuous.

      We also attempted to integrate mood into choice modelling. See Response 2 for Reviewer 3 for details.

      (4) The manuscript requires editing to improve clarity and precision. The use of terms such as "mood" and "approach motivation" is often inaccurate or not sufficiently specific. There are also many grammatical errors throughout the text.

      Thank you for this important suggestion. We have now explained motivation and mood in the Introduction section and the computational modeling section. Please see our clarifications below:

      Pages 3-4:

      “A growing literature indeed shows that risky behavior can be far better explained after adding value-insensitive approach and avoidance components to prospect theory [18,19], that is by including a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. This class of models highlights the important role of value-insensitive motivational components in decision making in addition to risk attitude-driven valuation (e.g., loss/risk aversion) [20].”

      Page 5:

      “Although mood is thought to persist for hours, days, or even weeks [30–33], momentary mood, measured over the timescale in the laboratory setting, represents the accumulation of the impact of multiple events at the scale of minutes [30,32,34–38]. Momentary mood external validity is demonstrated e.g., through its association with depression symptoms [37]. Mood is different from emotions, which reflect immediate affective reactivity and is more transient (e.g., from surprise to fear) [31–33,39].”

      We have corrected grammatical errors throughout the manuscript.

      (5) Claims of clinical relevance should be toned down, given that the findings are based on noisy parameter estimates whose clinical utility for the treatment of an individual patient is doubtful at best.

      Thank you for this comment. We agree that we did not evaluate the noise in our estimate e.g., by assessing the test-retest reliability on the task parameters, which is outside the scope of the study, and it is indeed possible that parameter estimate is somehow noisy. Therefore, we tone down the clinical relevance of our results. Please see our revision below:

      Page 32:

      “Next, we did not evaluate the noise in our estimate e.g., by assessing the test-retest reliability on the task parameters and it is indeed possible that parameter estimate is somehow noisy.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Title: I believe "aberrant mood dynamics" is both too general and overstating the results of this study, which did not measure mood dynamics longitudinally. "Aberrant" is also overly pathologizing. I would suggest sticking more directly to the results, for instance, "Insensitivity of momentary mood to non-risky rewards in adolescent suicidal patients".

      Thank you for this suggestion. We have now corrected it.

      (2) Abstract: in line 61, "Our study uncovers the cognitive and affective mechanisms" suggests that these are the only ones, and you uncovered them. Of course, there could be more mechanisms contributing to risk behavior in STB, so I would suggest removing the word "the" or adding "one of the".

      Thank you for this suggestion. We have now corrected it.

      (3) One major weakness of this study is that suicidal thoughts and behaviors were not assessed via a clinical instrument such as the Columbia Suicide Severity Rating Scale - this should be mentioned upfront.

      Thank you for this comment. According to medical records and information from family and friends by the researcher and psychiatrists, patients with suicidal thoughts and behaviors were categorized as suicidal group (S<sup>+</sup>), while patients without suicidal thoughts and behaviors were identified as control group (S<sup>-</sup>). Note that medical records and information were recorded from clinical interviews where the psychiatrists were vigilant for signs of suicidal ideation and inquired about suicidal-related thoughts and behaviors from both the patients and their families. Therefore, the current group operation was possibly comparable to Columbia Suicide Severity Rating Scale.

      (4) Table 1: female/male are sex, not gender (gender is man/woman/transgender/non-binary).

      Thank you for this suggestion. We have now corrected it.

      (5) Equation 1: It would be good to clarify what happens in gain-only or loss-only trials (the other value is then 0, but this can be clarified as it is not technically a loss or a gain).

      Thank you for this suggestion. We have now corrected it. Please see below for our revision:

      Page 12:

      “Please note that V<sub>gain</sub> is 0 in gain trials and V<sub>loss</sub> is 0 in loss trials.”

      (6) Figure 1E: The model prediction is not informative here. Given the linear regression model, there is no other option except that the mean prediction would overlap with the mean empirical measurement (unless the model was specified incorrectly). The same is true in Figure 2A.

      Thank you for this suggestion. We have now removed plots for model prediction.

      (7) Figure 1G: There was no analysis of the differences between groups in terms of earnings, given that the ANOVA was not significant. Still, if the claim is that risky behavior is sometimes suboptimal in this task, it would be good to show that there is a correlation between, say, symptoms of STB across groups and 1) risky behavior and 2) earnings.

      Thank you for this insightful comment. In the patient cohort, risky behavior (gambling rate)—but not earnings predicted the current suicidal ideation score (BSI-C, β = 9.189, t = 2.004, p = 0.048; earnings, β = 0.001, t = 0.582, p = 0.562). The lack of association for earnings is consistent with the task design, in which there is no stable optimal policy and payouts are only a coarse proxy for decision quality. Future work in learning paradigms, where optimality is well defined, may be better suited to test earning-based links to STB. We have clarified this point below:

      Page 32:

      “Second, although we assumed that increased risky behavior in STB was suboptimal, the current task was not suited to test this, given the task design of random feedback for gambling option. Future work in learning paradigms, where optimality is well defined, may be better suited to test earnings-based links to STB.”

      (8) Line 290: "beta_gain: -1-1" is unclear. I believe you meant beta_gain \in [-1,1].

      Thank you for this suggestion. We have now corrected it to make it clear.

      (9) The gain and loss biases are modeled as minimum and maximum probabilities for choosing the gamble. This is a legitimate choice for value-agnostic biases, but it is not the traditional choice (as far as I know). I wonder if the same results would hold with the more traditional formulation of the bias as an added constant to the utility of the gamble, i.e., p(gamble) = 1/(1+ exp(-mu(U_gamble + beta_gain - U_certain)). I believe in this case, you would also not have to specify different equations for positive or negative biases, or to limit the bias to the range of [-1,1] (indeed, the bias would be in reward-equivalent units).

      Thank you for this suggestion. The winning choice model we used here was consistent with previous literature (Rutledge et al., 2015 & 2016), which decomposed the decision process into risk-attitude-driven valuation (e.g., loss and risk aversion) and value-insensitive motivational components. These approach/avoidance parameters are a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference.

      As suggested, we also compared the traditional bias choice model. Model comparison did not support this. Please see Supplementary Page 4.

      (10) Also, for equations 5-8, it seems that 5-6 are identical to 7-8 except for the use of beta_gain versus beta_loss. You might want to consider simplifying by putting beta in the equations and specifying in the text that, depending on the trial type (loss or gain), the relevant beta is used.

      Thank you for this suggestion. We have now simplified it. Please see our revision below:

      (11) It is not clear what equations are applied to mixed trials in cM3.

      Sorry for the confusion. We have now clarified this point.

      Page 12:

      “Approach/avoidance parameters are not applied to in mixed trials.”

      (12) Model comparison: the mood models are nested within each other (e.g., mM3 can be derived from mM1 by setting beta_EV = beta_RPE). In this case, model comparison can use the likelihood ratio test instead of BIC, which can be too conservative (and therefore does not support the extra beta parameter for RPE, different from previous results in the literature). I wonder if a likelihood ratio test would lead to results more in line with previous findings with this task?

      Thanks for this suggestion. We agree that mM1 (CR+EV+RPE) and mM3 (CR+GR) are nested. However, our model space also included unnested models, such as mM5 (CR+GR<sub>better</sub>+GR<sub>worse</sub>). Therefore, it was not reasonable in our model space to use likelihood ratio tests.

      (13) Line 346: The replication sample is described as "healthy participants," however, their health (or mental health) status was not assessed, and they may as well have mental health concerns. I would suggest calling this a general sample or an undifferentiated sample - but not a healthy sample.

      Sorry for the confusion. We have now corrected this phrase.

      (14) Line 363: "in addition to the replication of previous findings in the validation dataset" is unclear. Are those tests not two-tailed?

      Sorry for the unclear statement. In the replication analyses, we used one-tailed t-tests because the direction of the effect was revealed on the clinical dataset. Please see our clarification below:

      Page 15:

      “For the replication of previous findings in the validation dataset, we used one-tailed tests in line with our clinically motivated directional hypothesis.”

      (15) Line 372: "validating our group manipulation" - the presented work does not have a manipulation. Maybe you meant "validating our grouping of participants"?

      Thank you for this suggestion. We have now corrected it to make it clear.

      (16) Figure 2B: It is not clear how the data were binned for illustration purposes only, and why this binning is necessary (I have not seen it in other papers) - presenting the data from each subject and the correlation line with error margins (as is done here) should be sufficient.

      Thank you for flagging this. For illustration only, we binned the data proportional to group sizes: in the patient sample (S<sup>-</sup> n = 25; S<sup>+</sup> n = 58; ≈1:2), we displayed 3 bins for S<sup>-</sup> and 6 bins for S<sup>+</sup>. We agree that binning is not necessary; all statistics were computed on raw, unbinned data. The binned panel was included solely for visualization, consistent with our prior work (Blain et al., 2023).

      (17) Table 2: delta BIC should be presented per subject (that is, divided by the number of subjects in each group), as the groups are of different sizes, so as presented now, the columns are not comparable across groups.

      Thank you for the helpful suggestion. Our goal in Table 2 is not to compare ΔBIC magnitudes across groups, but to identify the winning model within each group. The ΔBICs are aggregated at the group level solely to rank models for that group. Dividing by the number of participants would rescale each group’s column by a constant and would therefore not affect the within-group ranking or the conclusion that cM3 is the best model in all groups. For this reason, we retain the current presentation and interpret each column within group rather than across groups.

      (18) Line 640 - the effect of expectations and prediction errors on mood was not only shown in healthy people, but also in people with depression (Rutledge et al., 2007, https://pubmed.ncbi.nlm.nih.gov/28678984/)

      Thank you for this comment. Indeed, Rutledge et al., (2017) showed evidence for CR+EV+RPE mood model in adult people with depression. However, our study recruited adolescents with depression or anxiety, given that adolescent period might provide a developmental window for opportunities for early intervention of suicidality. Therefore, it is also possible that the current winning model was specific to adolescents. Please see our clarifications below:

      Page 28:

      “It is also possible that the current winning model was specific to adolescents. Given that Rutledge et al., (2017) supported the “CR-EV-RPE model” in adults with depression, our study with adolescent populations may suggest a developmental change for mood sensitivities.”

      (19) Supplemental material: Is the R2 section about R-squared? Perhaps you can use superscript on the 2 to make that clearer? For Figure S2, how was model recovery determined? Should I interpret the confusion matrix as suggesting that the winning model for each and every simulated subject was the generating model, or was the winning model determined for the whole simulated population in each of the 100 simulations? Traditionally, confusion matrices use the former measure, but the results of 100% recoverability make me suspect the latter was used here. In Figure S3, should we not be looking at simulated parameters and recovered parameters? What are "real parameters" here?

      Thank you for these important comments. We now consistently denote the coefficient of determination as R<sup>2</sup> (with a superscript 2) throughout the manuscript and Supplementary Materials.

      For the model recovery analysis in Figure S2, we have clarified that the confusion matrix is computed at the population level. Specifically, for each of the 100 simulations we generated a full dataset under each candidate model, fit all models to that dataset, and selected the winning model based on group-level model evidence (BIC). Each cell in the confusion matrix therefore reflects the proportion of simulations in which model j was selected as the best-fitting model when the data were generated by model i. This operation was reasonable because the decision of the winning model is made on the population-level dataset rather than on individual subjects.

      In Figure S3, the term “real parameters” referred to the parameters used to generate the simulated data. To avoid confusion, we now relabel these as “simulated (generating) parameters” and explicitly describe the figure as showing the relationship between simulated (generating) parameters and recovered parameters. Please see Supplementary Pages 2-3:

      “Model recovery: We generated 100 simulated datasets for each model (3 choice models and 8 mood models) using the fitted parameters of each model as the ground truth. Each dataset contained 201 trials and included 3 (or 8) sets of simulated data corresponding to the respective models. For each simulated dataset, we then fit all models and determined the winning model at the population level based on group-level BIC, yielding a confusion matrix in which each entry represents the proportion of simulations in which model j was selected as the best-fitting model when the data were generated by model i. As shown in Figure S2, all models are highly identifiable, indicating excellent recovery performance for both the choice and mood models.”

      “Parameter recovery: Figure S3 shows good parameter recovery for both choice and mood winning model (choice: rs > 0.91, ps < 0.001; intraclass coefficients > 0.78; mood: rs > 0.90, ps < 0.001; intraclass coefficients > 0.86). Moreover, we computed cross-correlations between all generating (“generating”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery.”

      Typos:

      (1) Line 90: original → originate

      (2) Line 596-598 - the same phrase is repeated twice.

      (3) Line 616: on the other word → hand.

      Sorry for the mistakes. We have now corrected them throughout the manuscript.

      Reviewer #2 (Recommendations for the authors):

      For people unfamiliar with interpersonal theory or motivational-volitional model, or three-step theory (lines 105-106), could you briefly explain the key idea of mood and suicide before going to the decision-making tasks? And from this, maybe motivate the predictions in your task? In particular, in the abstract and introduction, the phrasing could be a bit more concise and simpler. In the abstract, sentences were sometimes quite long. In the introduction, some paragraphs are somewhat repetitive. In the discussion, there were some typos.

      Thank you for these suggestions. We have now explained the key idea of mood and suicide before going to the decision-making tasks in the introduction, which can be seen below:

      Pages 4-5:

      “Contemporary theories of suicide converge on the idea that STB is initially caused by low mood experience. The interpersonal theory of suicide proposes that suicidal desire arises when people simultaneously feel socially disconnected (“thwarted belongingness”) and like a burden on others (“perceived burdensomeness”), experiences that are tightly linked to chronically low mood [25]. The motivational–volitional model [26] and the three-step theory [27,28] similarly emphasize that when negative mood and feelings of defeat or entrapment are experienced as inescapable, they can give rise to suicidal ideation, and that the progression from ideation to suicide attempts depends on additional factors such as reduced fear of death, increased pain tolerance, and a tendency to act impulsively under intense affect. Some official organizations, e.g., National Institute of Mental Health, have also listed mood problems as warning signals [8]. Interestingly, within the framework of decision making under uncertainty, gambling on lotteries with a revealed outcome has been found to induce high mood variance [29], providing an opportunity to assess the relationship between deficient mood and increased gambling decisions in STB.”

      We have also refined the wording and corrected typos throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Since many readers might only read the abstract, it is important that it is both informative and accurate. I have two suggestions in this respect. First, for the abstract to be more informative, it may be helpful to indicate already there that these are value-insensitive approach-avoidance parameters, in the sense that they favor/disfavor the gamble regardless of the potential outcomes' magnitude or probability. This issue is also present throughout the text, where the phrases "approach and avoidance motivation" are referred to as if they have established and precise computational definitions. In my view, these terms could just as easily be interpreted as parameters that multiply the value of potential gains or losses, which is not what the authors mean. It would be helpful to clarify this terminology.

      Thank you for these suggestions. In line with previous literature (Rutledge et al., 2015 & 2016), approach and avoidance motivation are indeed defined at the computational level, referring to a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. We have cited these papers in the manuscript. We also make it clear to further clarify approach and avoidance parameters in the abstract and introduction. Please see our revisions below:

      Page 2 (Abstract):

      “Using a prospect theory model enhanced with value-insensitive approach-avoidance parameters revealed that this rise in risky behavior resulted only from a heightened approach parameter in S<sup>+</sup>.”

      “Altogether, model-based choice data analysis indicated dysfunction in the approach system in S<sup>+</sup>, leading to greater propensity for gambling in the gain domain regardless of the lottery expected value.”

      Page 3 (Introduction):

      “A growing literature indeed shows that risky behavior can be far better explained after adding value-insensitive approach and avoidance components to prospect theory [18,19], that is by including a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. This class of models highlights the important role of value-insensitive motivational components in decision making in addition to risk attitude-driven valuation (e.g., loss/risk aversion) [20].”

      (2) The statement "our study uncovers the cognitive and affective mechanisms contributing to increased risk behavior in STB" is overstating the findings, as the study may have uncovered some contributing mechanisms, but likely not all of them. Removing the word "the" would fix this issue.

      Thank you for this suggestion. We have now corrected it.

      (3) Since mood is typically defined as lasting hours, it's inappropriate to refer to ratings that only reflect the last few trials as self-reports of mood. To be sure, I view the distinction between emotions and moods as quantitative, not qualitative, so I do not think there is a problem studying the former to understand the latter, but to avoid confusion, the terminology should follow common usage.

      Thank you for this suggestion. We follow previous work and operational definitions regarding mood (Rutledge et al., 2014, Eldar & Niv, 2015, Vinckier et al., 2018). Emotion is usually a very brief response to a specific stimulus (Emanuel & Eldar, 2023), e.g., leading to rapid changes like surprise then fear. In contrast, mood is defined as a diffuse state that is not specific to one stimulus. Here, we operationally and computationally define mood as an affective state reflecting the recent history of safe and gamble outcomes. We now clarify that point in the main text. Please see our revision below:

      Page 5:

      “Although mood is thought to persist for hours, days, or even weeks [30–33], momentary mood, measured over the timescale in the laboratory setting, represents the accumulation of the impact of multiple events at the scale of minutes [30,32,34–38]. Momentary mood external validity is demonstrated e.g., through its association with depression symptoms [37]. Mood is different from emotions, which reflect immediate affective reactivity and is more transient (e.g. from surprise to fear) [31–33,39].”

      (4) Line 78: The phrases "increase in risk attitude", "decrease in loss attitude", and "decrease in value-independent choice biases" are unclear to me in terms of their directionality. An attitude might be avoidant or embracing. If it is the former then increasing it would decrease risk-taking.

      Thank you for pointing out the ambiguity. We have now corrected them throughout the manuscript. Please see our revision below:

      Page 4:

      “We therefore hypothesized that heightened approach motivation, or weakened avoidance motivation, would account for increased risk behavior in STB.”

      (5) Line 125: I was not sure why one would expect the mood response to gamble-related quantities (EV and RPE) to be lower in STB and not higher.

      Sorry for the typo. We hypothesized that mood would respond more strongly to gambling-related quantities expected value (EV) and reward prediction error (RPE)—in adolescents with STB than in controls, given prior evidence that STB is associated with greater risk-taking.

      (6) The text could use proofreading, as there are many typos. These are from the first 100 lines alone:

      (a) Abstract: regardless the lotteries -> regardless of the lotteries'.

      (b) Line 78: it remains whether.

      (c) Line 80: can each -> each can.

      (d) Line 90: may original from.

      Sorry for the mistakes. We have now corrected them throughout the manuscript.

      (7) The rationale for focusing on the S+ group for mood model comparison is incorrect. The purpose is to identify parameters that vary as a function of suicidality, and for that, the S- group is just as important.

      Thank you for this comment. We agree that the S<sup>-</sup> group is as important as the S<sup>+</sup> group. A direct comparison was complicated because the winning mood models differed (S<sup>+</sup>: mM3; S<sup>-</sup>: mM5; Table 3). To ensure comparability, we checked results from both model specifications (mM3 and mM5). The conclusions were convergent: mood sensitivity to certain rewards (CR) was lower in S<sup>+</sup> than in S<sup>-</sup> (see Fig. 3 for mM3 and Fig. S8 for mM5).

      (8) There appears to be a contradiction between the inclusion criteria, which include having experienced suicidal thoughts and behaviors, and the definition of the S- group as not having suicidality.

      Thank you for pointing out this mistake. The corrected version of inclusion criteria can be seen on Page 7:

      “Patients were included if they met the following criteria: 1) both the researcher and psychiatrists agreed on their group classification; 2) they had a current diagnosis of major depressive disorder (MDD; unipolar depression), generalized anxiety disorder (GAD), or bipolar disorder with depressive episodes (BD), confirmed by two experienced psychiatrists using the Structured Clinical Interview for DSM-IV-TR-Patient Edition (SCID-P, 2/2001 revision; see Supplementary Note 1 for details);3) they were between 10 and 19 years of age; 4) they had no organic brain disorders, intellectual disability, or head trauma; 5) they had no history of substance abuse; 6) they had no experience of electroconvulsive therapy.”

      (9) It would be helpful to specify whether mood modeling was based on objective or subjective values, and why.

      Thank you for this helpful suggestion. We have now clarified whether mood modeling was based on objective or subjective values, and why. Specifically, we constructed two model families: one in which mood was driven by objective monetary outcomes (objective values) and one in which mood was driven by subjective values derived from each participant’s fitted choice model (subjective values). We then used the VBA_groupBMC function in the VBA toolbox to perform family-wise model comparison, with 8 candidate mood models within each family. Consistent with previous literature, the objective-value family provided a clearly superior fit to the data (exceedance probability, EP = 1.000). Based on this result and for parsimony, we report and interpret the mood modeling results from the objective-value family in the main text. We have clarified this point in Supplementary Note 9.

    1. eLife Assessment

      In this important study, the authors present an interesting platform for digital twin construction of iPSC-CMs using an AI-based approach. The concept is timely and could have meaningful impact as the field continues to explore integration of computational and experimental models. The evidence is convincing overall, although additional attention to framing and calibration of claims would enhance clarity and better reflect the current level of validation.

    2. Reviewer #2 (Public review):

      Summary:

      The authors present a computational framework for generating "cell-specific" digital twins of human iPSC-CMs from a single optimized voltage clamp recording. Using deep learning trained on > 1 million artificial cells, the authors demonstrate that the model can infer 52 biophysical parameters governing 6 major ionic currents, and the resulting digital twins can reproduce experimentally recorded action potentials.

      Comments on revised version:

      The authors propose an interesting platform for digital twin construction of iPSC-CMs using an AI-based approach. However, regarding the fundamental concerns raised in the previous review round "lack of experimental validation" and "overstatement of the claims", the authors have merely added text to the "Limitations" in the Discussion, without providing any new wet-lab experimental data. This cosmetic revision fails to demonstrate the scientific validity of the platform, and the core issues remain completely unresolved.

      I think the authors need to either provide substantial additional experimental data or drastically tone down the claims throughout the manuscript based on the following three major concerns.

      (1) Lack of wet validation

      The authors show that their AI model can infer 52 parameters from a single patch-clamp recording and reproduce the overall action potential waveform. However, the most critical validation (whether the individual ion channel parameters, such as IKr/ICaL, inferred by the AI actually match the true parameters of that specific cell) is still missing. Without a direct head-to-head comparison between the parameters inferred by the model and the exact values measured using conventional wet experiments, it is impossible to determine whether the platform is providing accurate prediction (or merely performing a curve-fitting).

      (2) Absence of experimental validation for drug response simulations (Cell 1 vs. Cell 2)

      In Figure 6, the authors present a simulation result where the administration of an IKr blocker (E-4031) induces EADs in the digital twin of Cell 1, but not Cell 2. However, there is absolutely no wet-lab validation for this prediction. Unless the authors actually administer the same drug to the live Cell 1 and Cell 2 from which the recordings were taken, this "computational drug response prediction" remains purely hypothetical. There is no evidence provided that the prediction accurately reflects real biological responses.

      (3) Significant overstatement regarding "inter-individual variability" and "personalized medicine"

      The authors state in the very first sentence of the Abstract: "Individual variability shapes how diseases manifest, how patients respond to therapy, and how rare phenotypes arise". However, this opening sentence is severely disconnected from the actual conclusions and data presented in this study. The platform can capture only "cell-to-cell variability within the same dish" (which is not even validated), and thus claiming "patient-to-patient differences" is an overstatement.

    3. Reviewer #3 (Public review):

      Summary:

      This work use convolution neural network to optimize a voltage clamp protocol to identify features and parameters from human pluripotent stem cell-derived cardiomyocytes.

      Strengths:

      The major strength is the methodology used to bridge in silico prediction of cell behavior and mechanistic insights from experimental dataset.

      Comments on revised version.

      As highlighted by the authors, due to the variability of the hPSC-CM model, to increase the applicability of this method, additional experimental dataset from different hPSC-CM lines would increase the translation of this approach.

      I personally found that the detailed description of the methods, including the rationale of including/excluding some parameters, is extremely helpful to whoever would like to use this approach in their research.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents an interesting approach for finding electrophysiological models that match experimental patch-clamp data. The authors develop a new method for deriving optimized current clamp protocols by training a neural network on synthetic data. This optimized current clamp is then used on both computational training data and on experimental data to predict current gating and conductance parameters that correctly reconstruct the electrical phenotype.

      Strengths:

      (1) The fitting of gating variables through an optimized patch clamp protocol is interesting.

      (2) The inclusion of experimental data is important, and the approach is shown to be effective in fitting them.

      Weaknesses:

      (1) Some clarity is necessary on the generation and selection of variable IPSC models. With such a large variation in so many parameters, I would expect some resulting parameters to generate non-realistic phenotypes, quiescent cells, etc. Are all 200,000 or 1,100,000 generated cells viable? Or are they selected somehow for realistic cell properties?

      Thank you for this important point. We agree that broad parameter variation can generate non-physiological model behavior. Indeed, with the +/-40% perturbation range, some simulated cells produced non-realistic outputs, including quiescent behavior, and failure to generate a complete action potential. These cases were excluded from the dataset. As a result, only cells exhibiting physiologically meaningful and numerically stable behavior were retained for further analysis. We have clarified this selection procedure in the Methods section. We applied a large variation to ensure that all possible combinations and morphologies were included in the training and testing data so the model would readily ingest new data and perform robustly.

      (2) The error shown in Figure 4 between different population sizes is not completely explained in the text - there seems to be a minimal difference between a population of 1,000 and 10,000, followed by a very good fit at 200,000. Is there a particular threshold that needs to be crossed where the error drops off? Related, how was the 200,000 number chosen?

      Thank you for this observation. We agree that the decrease in error shows a gradual performance improvement as the population size increases, rather than a strict cutoff. As shown in Figure 4, the difference between 1,000 and 10,000 samples is small, but as we continue to increase and get to around 200,000 samples, we see strong error minimization. This indicates how much training data is needed for optimal model performance. This improvement is due to better coverage of the high-dimensional parameter space, which helps the network learn the nonlinear relationships between the parameters and outputs.

      We tested a range of training data sets and found that above 200,000 training data sets, the model consistently produced low, stable errors and good test-training agreement. The test error decreased with the training error as the population size increased, indicating better generalization and suggesting that the model accurately predicts unseen data rather than overfitting to the training set.

      (3) Related to the point above, the 1,100,000 population for fitting experimental data also needs a more complete explanation: how was this number chosen, and how does the error compare with the other population sizes shown in Figure 4?

      Thank you for this question. We found that at a training data set size of 1,100,000 we were able to cover the large parameter space induced by +/-40% parameter perturbation. iPSC-CM measurements are known to exhibit high variability, and we wanted to capture the full range in the training data set so the model could ingest a wide range of experimental data. It is trivial to generate new training data, for example, to capture different experimental conditions like temperature differences, mutations, drugs, or ionic variability. We view this flexibility as a substantial strength of the approach. But the large perturbations we show in this study (+/-40%) allow the generation of a very broad range of cellular phenotypes while maintaining physiologically realistic ionic current properties and action potential behavior. Consistent with Figure 4, increasing population size reduces prediction error and improves generalization. The larger dataset provided more stable, accurate predictions when fitting experimental data, without evidence of overfitting.

      (4) Why are the optimized current clamp protocols different between panels A and B in Figure 5? Are they somehow informed by experimental data?

      Thank you for this question. The stimulation protocol used in panels A and B is identical. Panels A and B show whole-cell currents recorded under the same stimulation conditions as in Figure 3. The differences reflect variability in the underlying whole-cell ionic currents of the model cells rather than differences in the applied protocol. This is exactly the idea: the exact same protocol will generate different whole-cell currents in individual cells, but the model can find parameter sets for all of them.

      (5) Figure 6D: Is the EAD risk in panel D specific to cell 1, 2, or the pooled variants of both?

      Thank you for this question. We have clarified this point in the revised manuscript. The EAD risk shown in panel D is computed from the pooled variants of both Cell 1 and Cell 2, rather than being specific to either cell individually.

      (6) How sensitive is the fitting to minor parameter variation? Further, if one were to pick, let's say, the next-best-fitting value, would that fall close to the best one? Is the solution found unique, or are there multiple sets with good fits?

      Traditional optimization methods, such as Nelder–Mead, directly fit the model to the observed data by iteratively minimizing the error for each dataset. As a result, the solution can depend on the initial parameter guess and may converge to different local minima. In contrast, our approach trains a deep learning model on synthetic data generated from the baseline model, learning a mapping from whole-cell currents to the corresponding 52-parameter sets by minimizing prediction error. The mean squared error (MSE) decreases from approximately 10⁻² to below 10⁻³, with training and test errors overlapping closely, indicating stable training, good generalization, and accurate reproduction of the observed signals.

      The model achieves very low MSE and reproduces the electrophysiological outputs with high fidelity. However, accurate reproduction of the outputs does not imply a unique parameter solution. This is illustrated in Figure S1, where baseline and predicted parameter values show close agreement overall, yet small deviations persist across parameters. This indicates that different parameter combinations can yield similar whole-cell behaviors due to parameter correlations and compensatory effects. In such cases, the model learns to predict a representative parameter set that is most consistent with the training data and loss function, rather than converging to a single unique solution within a fixed numerical tolerance.

      Reviewer #2 (Public review):

      Summary:

      The authors present a computational framework for generating "cell-specific" digital twins of human iPSC-CMs from a single optimized voltage clamp recording. Using deep learning trained on > 1 million artificial cells, the authors demonstrate that the model can infer 52 biophysical parameters governing 6 major ionic currents, and the resulting digital twins can reproduce experimentally recorded action potentials.

      Strengths:

      The framework has clear potential for understanding cellular heterogeneity in iPSC-CMs, predicting individual drug responses, and reducing the experimental burden of multiple patch clamp protocols.

      Weaknesses:

      There are several concerns about the validation of the model and its clarity. First, the biological variability being modeled in this manuscript is not defined well. It is unclear whether the framework addresses cell-to-cell differences within a single differentiation batch, variability across iPSC lines, or donor-to-donor differences. This ambiguity makes it difficult to interpret what the "digital twin populations" actually represent biologically. Second, the main claim, "the digital twins enable drug testing and arrhythmia prediction that would be impractical experimentally", is not experimentally validated. For example, the E-4031 simulations predict EAD rates, but no direct experimental head-to-head comparison is provided to confirm that these predictions are accurate. Third, technical reproducibility and biological representativeness are not assessed. Single voltage clamp recordings are inherently noisy. Without knowing how much variability comes from the recording process (technical variation) vs true biological differences, it is difficult to judge whether observed "cell-specific" parameter differences are meaningful. In addition, the optimized protocol is claimed to be superior to conventional approaches, but again, no experimental comparison is shown.

      The authors should address these concerns, with particular emphasis on clarifying the biological context and providing direct experimental validation. Below are detailed specific points:

      (1) Ambiguous definition of iPSC-CM heterogeneity. The authors model "typical iPSC-CM heterogeneity" by varying 52 parameters +/- 40% around a baseline model (Figure 1), generating > 1 million synthetic cells. However, the manuscript does not clearly state what biological variability this model is intended to capture. Is this modeling within-line, cell-to-cell variability (e.g., cells from the same dish or differentiation batch that differ due to stochastic gene expression or maturation state)? Or is this modeling between-line or between-donor variability (e.g., genetic background differences, reprogramming efficiency)? This distinction is critical for interpretation. If the goal is to understand why different cells in the same dish behave differently, then training data should reflect that. If the goal is to compare patient lines or disease models, the framework needs validation across multiple donors or lines.

      For example, the experimental validation in Figure 5 uses a single iPSC line (iPS-6-9-9T.B), but how many differentiation batches or dishes were tested, or whether cells came from the same preparation are unclear. Another example is that the wide AP diversity in the training population (Figure 1A) is impressive, but there is no demonstration that real experimental cells actually fall within this assumption range of +/- 40%.

      From a biological perspective, iPSC-CMs are known to be highly heterogeneous within lines (maturation state, metabolic differences, epigenetic variation, spatial differences within the same dish, etc) and between lines (different donor/genetic background). Thus, please explicitly state whether the +/- 40% variation is intended to model within-line or between-line heterogeneity, and justify this choice with wet experiment data (or reference to experimental literature on iPSC-CM variability). Please clarify how many dishes, differentiation batches, and time points post-differentiation were used for experimental recordings (Figures 5-6). If the framework is intended to generalize across lines from different donors, please test the model on multiple independent iPSC lines (from different donors).

      Thank you for this important and insightful comment. The selected ±40% range was chosen to broadly explore all physiologically plausible electrophysiological behaviors, not to match a specific experimental distribution. Our goal was to cover enough behaviors for the model to learn a reliable mapping between responses and ionic parameters.

      We recognize that this approach does not explicitly account for variability between lines or donors. We have a current project focused on extending the framework to include multiple iPSC-CMs from patient donors, but given that the model framework successfully reproduces such a broad range of cell phenotypes, we feel confident that it will readily apply to different genetic backgrounds from patient-specific cells. This study is underway.

      We have updated the manuscript to clarify how the modeled variability is interpreted and added a discussion of these limitations. Furthermore, we clarified the experimental conditions, such as the number of differentiation batches and recording settings, in the revised Methods section.

      (2) Biological representativeness of single-cell measurements.

      The framework generates digital twins from single voltage clamp recordings. The patch clamp recordings in iPSC-CMs are subject to substantial technical variability. The manuscript does not address a fundamental question: "How representative are the measurements from a single cell on the dish (or line)?" In other words, if I measure one cell from a dish of a million cells, does that cell's digital twin tell me something about the dish as a whole, or just about that one cell? The manuscript presents Cell 1 and Cell 2 (Figures 5-6) as distinct individuals, but it's unclear whether these differences reflect true biological heterogeneity or simply sampling variability. I think the authors should perform replicate recordings on multiple cells (e.g., > 10 cells) from the same dish (same differentiation batch) and quantify how much the inferred parameters vary, and then compare between lines.

      Thank you for this important comment. We agree that the representativeness of single-cell measurements and the impact of technical variability are important considerations in interpreting the results. In this study, the framework is designed to generate digital twins that reflect the electrophysiological properties of individual recorded cells, rather than to directly represent the behavior of the entire cell population within a dish.

      As such, differences observed between Cell 1 and Cell 2 are intended to reflect variability at the single-cell level, which may arise from a combination of biological heterogeneity and experimental variability. We agree that systematic replicate recordings across multiple cells are valuable to quantify the relative contributions of biological and technical variability, and to assess the consistency of inferred parameters. However, this is beyond the scope of the current study. We have added clarification in the manuscript to explicitly state this limitation and to outline this as an important direction for future work.

      (3) No experimental validation of the main claim that in silico populations can replace wet experiments.

      The most exciting claim in the manuscript is that digital twins enable drug testing and arrhythmia prediction "at scale" without requiring hundreds of patch clamp experiments. Specifically, the authors show that in silico populations derived from two experimental cells (Figure 6C) predict dose-dependent EAD incidence for the IKr blocker E-4031 (Figure 6D), with ~3% of cells showing EADs at 50 nM.

      However, this prediction is not validated experimentally. If I actually patch 20-30 real iPSC-CMs and apply 50 nM E-4031, will ~3% of them show EADs, as the model predicts? Without this validation, I think the drug testing framework is purely hypothetical. The model may be internally consistent (e.g., Cell 1's twin behaves differently from Cell 2's twin), but there is no evidence that these in silico populations reflect real biological variability in drug response. Please provide experimental validation that justifies the prediction by digital twins.

      Thank you for this important comment. We agree that experimental validation of population-level drug response will be valuable for establishing the quantitative accuracy of the predicted EAD incidence. The E-4031 simulations are intended as a proof-of-concept illustrating how the framework can identify susceptible subpopulations and quantify relative proarrhythmic risk in silico. We agree that direct comparison with large-scale experimental datasets is a key next step, and we are working hard to get the study funded so that we can perform those experiments and bring this technology to scale.

      (4) Experimental validation and head-to-head comparison of optimized protocol.

      The authors claim that their deep learning-optimized voltage clamp protocol (Figure 3, Figure 4A) is superior to conventional approaches, but they have not validated this experimentally by doing a head-to-head comparison. The manuscript does not compare the optimized protocol to any published voltage clamp designs. If the optimized protocol is genuinely easier to implement and more informative than existing approaches, this would be a major practical advance. But without side-by-side comparison, it is impossible to judge whether the optimization made a real difference.

      Thank you for your comment. We agree that comparing directly with traditional voltage-clamp protocols through experiments would be useful. In this study, our main aim was to show that the optimized protocol enhances parameter inference within the modeling framework, not to prove experimental superiority. We have clarified this point in the revised version.

      Reviewer #3 (Public review):

      Summary:

      This work uses a convolutional neural network to optimize a voltage clamp protocol to identify features and parameters from human pluripotent stem cell-derived cardiomyocytes.

      Yang et al. introduce an innovative experimental framework that integrates computational modeling and deep learning to generate a digital twin of human pluripotent stem cell-derived cardiomyocytes (hPSC-CMs).

      Strengths:

      The major strength is the methodology used to bridge in silico prediction of cell behavior and mechanistic insights from the experimental dataset.

      The approach used in this study represents a significant step toward precision medicine by enabling in silico prediction of cellular behavior and mechanistic insight from experimental datasets. The study addresses an important and timely challenge in stem cell-based and personalized medicine, and the authors compellingly leverage state-of-the-art methods alongside strong expertise in computational modeling and cardiac electrophysiology

      Weaknesses:

      While the overall approach is highly compelling and the potential impact is substantial, there are two areas where clarification and refinement, particularly in the phrasing and framing used throughout the manuscript, would further strengthen the work.

      (1) While the overall goal of the study is compelling, the manuscript would benefit from clearer articulation of how the proposed framework is intended to be used in practice. In particular, it is not entirely clear whether the authors envision this approach as:

      (a) a method to extract population-level trends that, when paired with biological data, enhance statistical power and interpretability, or

      (b) a strategy capable of constructing a population-based model from limited single-cell recordings. If the latter is intended, additional guidance on the number of action potentials required per cell and the assumptions underlying this extrapolation would greatly clarify the scope and applicability of the method.

      Thank you for this thoughtful comment. We agree that the intended use of the framework should be more clearly articulated. In this study, we generate a large synthetic population of iPSC-CM models by varying 52 biophysical parameters governing key ionic currents. A neural network is trained on simulated whole-cell current responses to learn a mapping between current profiles and model parameters. Experimental recordings are then used as inputs to this trained model to infer ionic parameters, rather than directly fitting the model to data. This enables individual recordings to be interpreted within a large, physiologically plausible parameter space and supports population-level analysis of electrophysiological variability. The primary goal of the framework is therefore to facilitate mechanistic interpretation of variability and relate experimental observations to underlying ionic currents. But the longer-term intended goal is to develop digital twins from patient-derived cell lines and then use populations constructed from patient-specific digital twins to screen therapeutics and identify arrhythmia marker vulnerability in a very thorough and high-throughput way. We have clarified this in the revised manuscript.

      (2) The manuscript would also benefit from a clearer explanation of how electrophysiological heterogeneity observed in hPSC-CMs is linked to inter-patient variability. Although the authors state that this framework can be generalized to compare patient-specific hiPSC-CM lines, it remains unclear how this generalization is achieved, given the substantial sources of variability intrinsic to hiPSC-CMs (e.g., batch effects, reprogramming strategy, differentiation protocol, and maturation state). As acknowledged by the authors, addressing this level of variability likely requires large datasets; further clarification of how the proposed approach mitigates or accommodates these challenges would strengthen the translational claims.

      Below are my suggestions that could help strengthen the claims in the manuscript:

      (1) Adding a dedicated section describing the electrophysiological phenotype of the hPSC-CMs used in this study would help justify the choice of the underlying ionic model and the selection of the six ion currents analyzed. These currents are not only developmentally regulated but may also vary substantially across different hPSC-CM lines, which has implications for generalizability.

      Thank you for this important suggestion. We agree that providing additional context on the electrophysiological phenotype of the hPSC-CMs strengthens the rationale for both the underlying ionic model and the selection of currents analyzed.

      We have expanded the Methods section to clarify this point. Briefly, the ionic currents were selected based on the Kernik-Clancy iPSC-CM model developed in our prior work, which was specifically designed to capture the range of electrophysiological variability observed within an iPSC-CM cell line using a population-based framework. In this model, variation in key ionic conductances is sufficient to reproduce the diversity of action potential morphologies, spontaneous activity, and repolarization dynamics commonly reported experimentally, while avoiding non-physiological behaviors.

      Accordingly, we focused on six primary ionic currents that are known to play dominant roles in shaping action potential characteristics and variability in iPSC-CMs. This selection reflects a balance between model parsimony and physiological relevance, enabling the framework to capture the expected spectrum of variability within a given cell line. We also note that the framework is extensible, and additional currents or alternative parameterizations can be incorporated to account for differences across cell lines, donors, or experimental conditions in future studies. See updated discussion.

      (2) If feasible, inclusion of patch-clamp data from an additional hPSC-CM line would significantly strengthen the claim that this framework can harmonize and generalize across datasets and cell sources.

      Thank you for this helpful suggestion. We agree that adding data from more hPSC-CM lines would improve the framework's generalizability. In this work, our goal was to show that the digital twin framework is data-driven and can easily be expanded to include more hPSC-CM lines, allowing for cross-line comparisons in future studies. We have clarified this and included a discussion of this limitation in the revised manuscript. We are currently seeking funding for patient-specific lines as well to allow scalability.

      (3) The authors note that the experimental cells exhibited high variability in action potential morphology. This is an important observation that directly supports the motivation for the study and should be explicitly presented, even if only in the supplementary materials.

      Thank you for this suggestion. We agree that explicitly showing the variability in experimental action potential morphology strengthens the motivation for this study. We have now added a section in the discussion discussing this and referencing the many prior studies that focused on iPSC-CM variability, including the studies upon which our initial model (Kernik-Clancy) was based.

      (4) In the hERG-blocker experiments, further clarification is needed regarding the biological relevance of the reported 3% incidence of early after depolarizations (EADs). Additionally, an interrupted sentence in this section makes it unclear whether the goal is to demonstrate that the digital twin can capture rare arrhythmic risk events or whether the digital twin is necessary to determine whether this level of risk is clinically meaningful.

      Thank you for this important comment. We agree that more clarification is needed on the ~3% EAD incidence and the digital-twin role. This analysis aims to show that electrophysiological variability can create a small, susceptible subpopulation under drug effects, not to set a clinical risk threshold. The observed ~3% EAD incidence reflects the emergence of such a susceptible subpopulation under hERG block. While relatively small, this fraction is important because it arises from modest, physiologically plausible variation in ionic properties and would be difficult to capture using single-cell or small-sample approaches. As described in the Discussion, this variability-driven emergence of EADs provides a quantitative measure of proarrhythmic risk at the population level. The digital-twin framework enables systematic identification and quantification of these rare events, linking cell-level variability to population-level responses. We have revised the manuscript to clarify this point.

      (5) The manuscript states that some action potentials were excluded from the experimental dataset. A brief explanation of the exclusion criteria, along with guidance on how to distinguish high-quality from low-quality recordings, would improve transparency and reproducibility.

      Thank you for this comment. We agree that the definition of failed recordings should be clarified. We have now specified the exclusion criteria in the Methods section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It would be helpful if the network cartoon in Figures 2 and 3 were replaced with a simplified sketch of the actual neural network used.

      Thank you. We now have new figures 2 and 3.

      (2) Subsection title for the Introduction has a typo.

      Thank you. We have fixed it.

      Reviewer #2 (Recommendations for the authors):

      (1) Technical quality control criteria are not specified.

      The Methods section states that "any incomplete or failed recordings were excluded," but does not define what constitutes a failed recording. The criteria could be subjective.

      Thank you for pointing this out. We agree that the definition of failed recordings should be clarified. We have now specified the exclusion criteria in the Methods section.

      “Recordings were excluded if they exhibited no spontaneous firing, abnormally slow firing rates, or failed to capture a complete action potential waveform. These criteria were applied consistently across all recordings.”

      (2) "Cell-specific" may overstate the claim.

      The term "cell-specific digital twins" (title, throughout) implies that the inferred parameters reflect the true biological state of each cell. However, parameters are derived only from curve-fitting to electrophysiological data and do not reflect other biological components (e.g., gene expression, contractility, calcium handling, metabolism, etc). Please consider rephrasing to "electrophysiology-based digital twins", "voltage clamp-matched digital twins", etc.

      Thank you for this important comment. We agree that the term “cell-specific” could be interpreted as implying a complete representation of the biological state of each cell. We have also adjusted the wording in relevant sections to avoid over-interpretation.

      Reviewer #3 (Recommendations for the authors):

      (1) I would add the list of the 52 parameters in the method section/SI and not just in the reference. Additional justification of why the perturbation was set as +/- 40% for the 52 parameter or +/- 20% for the EAD population would also help.

      Thank you for this helpful comment. We have included model equations and highlighted the 52 parameters in the Supplementary Information and provided additional justification in the Methods.

      (2) In Figure 1B, might be helpful to add the axis of the Vm instead of the dotted line indicating 0 mV to show differences in the diastolic potential.

      Thank you! We have now updated Figure 1B.

      (3) Figure 1C-I might be more impactful to show traces from the AP shown in Figure B to reinforce the impact of a single current in the AP shape.

      We have now updated Figure 1C-I to include traces from the AP shown in Figure 1B.

    1. eLife Assessment

      This important study shows that long-range somatostatin-expressing neurons in the ventrolateral periaqueductal grey that project to the rostral ventromedial medulla selectively suppress pain responses during conditioned fear. The evidence supporting these conclusions is exceptional, with methods spanning a novel cued fear-conditioned analgesia paradigm, cell-type-specific optogenetic activation and inhibition, anatomical circuit tracing, and in vivo spinal cord electrophysiology. These results will be of broad interest to systems and behavioral neuroscientists studying fear, pain, and descending pain-control circuitry.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In the manuscript by Winke et al, the authors present evidence that fear-induced analgesia is mediated by somatostatin projection cells from the vlPAG to the RVM. This study uses a mouse model of fear-induced analgesia, and incorporates optogenetic circuit manipulation with behaviour and electrophysiology to gain a meaningful insight into a novel circuit involved in fear-induced analgesia.

      Strengths:

      (1) This is a well-constructed study with appropriate controls and analyses.

      (2) Alternative interpretations of the data are systematically considered and eliminated via rational experiments. The authors are commended for a nice piece of experimental work.

      (3) The vlPAG is a known region of pain modulation, and this study adds valuable insight to the circuit involved in fear-associated analgesia.

      Weaknesses:

      Only male mice are included in this study. [This has been explained and noted as a limitation.]

    3. Reviewer #2 (Public review):

      Summary:

      Wenke et al. investigated the role of vlPAG somatostatin-expressing neurons in the mediation of analgesia during defensive states. A newly developed paradigm of cued fear-conditioned analgesia, which consists of a combination of an auditory fear retrieval session and a pain test, was used to evaluate this cell population's contribution to fear-mediated analgesia. Optogenetic manipulation of vlPAG SST+ neurons modulated the responses to a nociceptive cue (Hot Plate) presented concomitantly with an aversively conditioned tone. At the same time, alterations in the freezing levels could be observed during optogenetic activation of vlPAG SST+ neurons. In order to disentangle the impact of these cells on analgesia from their impact on the expression of defensive behaviors, the authors performed electrophysiological recordings from the dorsal horn in the spinal cord of anesthetized mice. A vlPAG-RVM-DH pathway was identified to trigger nociceptive C-fibers upon optic activation of the RVM. Finally, pathway-specific activation of SST+ vlPAG-RVM neurons could abolish CS-induced analgesia.

      Strengths:

      The study addresses a relevant topic, that is, brainstem circuits for pain-modulatory mechanisms as part of defensive states evoked by threat. This is important because the circuit mechanisms underlying pain are still not fully understood, and defining molecular markers of cellular circuit substrates may support the identification of potential pharmaceutical targets in treating pain. The authors confirm a previous study in that a somatostatin-positive cellular population presents a crucial vlPAG circuit element mediating anti-nociceptive effects. Key novelty aspects of the present study are the demonstration that these neurons seem to play a role specifically in threat-induced analgesia. This was possible by the elegant design and application of a novel fear analgesia paradigm, combined with cell- and pathway-specific optogenetics.

    4. Reviewer #3 (Public review):

      Summary:

      Conditioned analgesia refers to the ability of a learned fear cue to suppress pain-related behavior and neural activity. Understudied, the authors developed a novel conditioned analgesia procedure in which a cue that had been paired or unpaired with shock was played while a hot plate increased temperature. Compared to several control conditions, the authors found increased latency to a nociceptive response (paw licking). The authors identified somatostatin neurons in the periaqueductal gray as a likely mediator of the behavior. They then showed that: (1) stimulating vlPAG-SST neurons blocked nociceptive response latency increases to the CS+, (2) stimulating vlPAG-SST neurons suppressed fear retrieval freezing, (3) stimulating vs. inhibiting vlPAG-SST neurons drove opposing modulation of c-fibers and Aδ-fibers, (4) direct-projecting vlPAG SST neurons modulate freezing while RVM-projecting vlPAG SST neurons modulate conditioned analgesia.

      Strengths:

      These experiments have many strengths. The behavioral assay is chief among them. The assay is robust and controls for confounding factors to reveal a repeatable effect of a shock-paired cue to delay nociceptive responding. The optogenetic experiments provide the correct level of temporal precision, given the authors' time-specific interest in cued responding. Combining neuronal manipulations with spinal recordings is particularly innovative, especially in the context of more behavioral neuroscience-based assays. All-in-all, I found this to be an exceptionally strong set of experiments.

      Weaknesses:

      No obvious weaknesses were identified by this reviewer.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript by Winke et al, the authors present evidence that fear-induced analgesia is mediated by somatostatin projection cells from the vlPAG to the RVM. This study uses a mouse model of fear-induced analgesia, and incorporates optogenetic circuit manipulation with behaviour and electrophysiology to gain a meaningful insight into a novel circuit involved in fear-induced analgesia.

      Strengths:

      (1) This is a well-constructed study with appropriate controls and analyses.

      (2) Alternative interpretations of the data are systematically considered and eliminated via rational experiments. The authors are commended for a nice piece of experimental work.

      (3) The vlPAG is a known region of pain modulation, and this study adds valuable insight to the circuit involved in fear-associated analgesia.

      We are very thankful to the referee for these positive comments.

      Weaknesses:

      (1) Only male mice are included in this study.

      We thank the reviewer for this point. We used only males in this first study for practical reasons to work with a population as homogeneous as possible. However, taking sex differences in biological mechanisms into account, we included this restriction in the summary and discussion

      (2) Animals are excluded from analyses based on clearly defined criteria, but it is not clear how many mice were excluded from each group.

      We thank the reviewers for raising this point. As stated in the Methods, we applied strict inclusion criteria for mice undergoing the hot-plate test, specifically a discrimination index ≥ 0.4 and a conditioning index ≥ 0.3. Using these criteria, 23% of wild-type mice were excluded for failing to meet the discrimination criterion. In the transgenic groups, an average of 20% of mice failed to meet the learning criteria, and an additional 12% were excluded due to incorrect opsin injection or misplaced optic fiber placement.

      (3) The authors implement a pain sensitivity assay that involves a hot plate with progressively increasing temperature. The time to nociceptive responses is reported. Without reporting the actual temperature at which the mice respond, it makes it difficult to compare nociceptive responses to previously published work (which typically use a defined and static hotplate temperature).

      We thank the reviewer for this comment. We provided this information related to the actual temperature of the nociceptive response in the original manuscript in supplementary figures 1, 2 and 5.

      (4) The authors present evidence that inhibition of SST vlPAG cells enhances spinal nociceptive electrophysiological responses, but the corresponding pain sensitivity is not altered (Figure 2, CS- condition). The reason for the discrepancy between electrophysiological and behavioural responses is not clear.

      We believe this comment arises from a misunderstanding of our results. In our study, inhibiting SST+ vlPAG cells did not increase nociceptive electrophysiological responses. Instead, it decreased spinal nociceptive transmission, as evidenced by reduced nociceptive field potentials and WDR responses in Figure 4c,e. Consistent with this electrophysiological effect, photoinhibition of SST+ vlPAG cells also produced behavioral analgesia, as evidenced by increased nociceptive response latency in the hotplate test under both CS− and CS+ conditions (Figure 2f). Therefore, our electrophysiological and behavioral findings are not contradictory but instead support the conclusion that inhibiting SST+ vlPAG cells reduces pain sensitivity regardless of defensive state. We will revise the text to clarify this point.

      Reviewer #2 (Public review):

      Summary:

      Wenke et al. investigated the role of vlPAG somatostatin-expressing neurons in the mediation of analgesia during defensive states. A newly developed paradigm of cued fear-conditioned analgesia, which consists of a combination of an auditory fear retrieval session and a pain test, was used to evaluate this cell population's contribution to fear-mediated analgesia. Optogenetic manipulation of vlPAG SST+ neurons modulated the responses to a nociceptive cue (Hot Plate) presented concomitantly with an aversively conditioned tone. At the same time, alterations in the freezing levels could be observed during optogenetic activation of vlPAG SST+ neurons. In order to disentangle the impact of these cells on analgesia from their impact on the expression of defensive behaviors, the authors performed electrophysiological recordings from the dorsal horn in the spinal cord of anesthetized mice. A vlPAG-RVM-DH pathway was identified to trigger nociceptive C-fibers upon optic activation of the RVM. Finally, pathway-specific activation of SST+ vlPAG-RVM neurons could abolish CS-induced analgesia.

      Strengths:

      The study addresses a relevant topic, that is, brainstem circuits for pain-modulatory mechanisms as part of defensive states evoked by threat. This is important because the circuit mechanisms underlying pain are still not fully understood, and defining molecular markers of cellular circuit substrates may support the identification of potential pharmaceutical targets in treating pain. The authors confirm a previous study in that a somatostatin-positive cellular population presents a crucial vlPAG circuit element mediating anti-nociceptive effects. Key novelty aspects of the present study are the demonstration that these neurons seem to play a role specifically in threat-induced analgesia. This was possible by the elegant design and application of a novel fear analgesia paradigm, combined with cell- and pathway specific optogenetics.

      We thank the referee for such positive feedback.

      Weaknesses:

      Despite the convincing and rigorous experimental approach, the study leaves some interpretational room when it comes to the proposed circuit mechanism. This could either be addressed by additional experiments or by more discussion of alternative circuit layouts.

      Major Comments:

      (1) The paper by Zhang et al. (https://pubmed.ncbi.nlm.nih.gov/36641028/), which identified a role for vlPAG SOM+ neurons in mediating anti-nociception in neuropathic pain, needs to be referenced and its results discussed, if not reconciled. While functionally, both studies find an analgetic role of vlPAG SOM+ neurons projecting to the RVM, Zhang et al., using slice physiology, characterize those neurons as glutamatergic. In Figure 4E of Zhang et al. they find general (fear-independent) analgetic effects with PAG-RVM specificity by performing chemogenetic experiments.

      We thank the reviewer for highlighting this important point. We agree that the study by Zhang et al. is highly relevant and should be discussed in the revised manuscript. Their work shows that inhibiting vlPAG SST/SOM neurons with chemogenetic methods produces analgesia in a neuropathic pain model, and in our study, we similarly found that inhibiting SST+ vlPAG neurons increases hotplate response latency (Figure 2f), which aligns with an analgesic effect. Additionally, we observed that activating SST+ vlPAG neurons suppresses fear-conditioned analgesia.

      At the same time, there are important differences between the two studies that may explain the differences in interpretation. First, the behavioral paradigms are not identical. Zhang et al. used a hotplate protocol where animals were directly exposed to a nociceptive temperature, whereas in our study, we used a progressive temperature ramp and explicitly compared responses during a conditioned stimulus (CS+) and a non-conditioned control stimulus (CS−). These controls were important for us to distinguish fear-specific effects from more general effects related to stress, arousal, sensitization, or other non-associative processes.

      Second, the two studies differ in experimental context. Zhang et al. examined this circuit in a neuropathic pain model, whereas our study focused on acute nociceptive processing and fear-conditioned modulation of pain. We therefore believe that the apparent discrepancy might reflect differences in pain state and behavioral context, rather than a direct contradiction.

      Finally, Zhang et al. showed in slice recordings that SST+ vlPAG neurons provide excitatory input to RVM neurons. This is an important finding that we now address in the revised manuscript. At the same time, because the RVM contains heterogeneous neuronal populations with different projection targets and functions, these recordings alone do not prove that all recorded RVM neurons are part of the descending pathway controlling spinal nociception. Therefore, we have revised the Discussion to explicitly acknowledge Zhang et al. and to emphasize both the similarities and differences between the two studies.

      It can be argued that in addition to the two functionally distinct inhibitory SOM subtypes hypothesized by Winke et al., there is another, excitatory subpopulation. Also, the different experimental conditions (chronic vs. acute pain, non-threat vs. fearful cues/contexts may recruit different vlPAG SOM+ populations. All of this is conceivable, yet I wonder whether the contrasting findings could more parsimoniously be reconciled. The author's own results presented here in Supplementary Figure 3 suggests that SOM+ vlPAG cells are colocalizing with glutamate and thus could also be excitatory. In addition to this rather complementary piece of evidence, a more extensive characterization of vlPAG neurons using IHC and slice physiology would be needed to justify the unambiguous identification of their inhibitory nature.

      We thank the reviewer for this thoughtful comment. We agree that our current data do not support a definitive conclusion that all SST+ vlPAG neurons are inhibitory. As the reviewer notes, our Supplementary Figure 3 shows that SST+ vlPAG cells can also co-localize with glutamatergic markers, which is consistent with the possibility of cellular heterogeneity within this population. We also agree that different experimental conditions, such as chronic versus acute pain and non-threatening versus fear-related contexts, may activate different SST+ vlPAG subpopulations.

      Our intention was not to claim that SST+ vlPAG neurons constitute a uniform inhibitory population, but rather that SST+ cells are strongly represented among inhibitory neurons in the vlPAG. We agree, however, that more detailed characterization, including additional immunohistochemical analyses and slice physiology, is necessary to more definitively determine the neurotransmitter phenotype and functional connectivity of these neurons. We have therefore revised the text to temper our interpretation and to explicitly acknowledge the likely heterogeneity of SST+ vlPAG neurons, including the possibility of an excitatory subpopulation. We therefore modified the discussion accordingly:

      “Our results align with the parallel inhibition- excitation model, where inhibitory and excitatory cells form two distinct, parallel descending pathways for pain modulation.

      Indeed, previous research demonstrated the presence of an inhibitory pathway projecting throughout the PAG–RVM-spinal cord dorsal horn neuraxis. Our results complement this study by suggesting that one of these previously proposed parallel pathways is mediated by SST+ vlPAG cells and has a functional role in mediating analgesia. At the same time, our data indicate that vlPAG SST neurons are heterogeneous, with approximately one-third of these cells co-localizing with excitatory markers. Together with the recent observation that excitatory SST+ vlPAG neurons project to the RVM (Zhang et al., 2023), this raises the possibility that a subset of long-range SST+ vlPAG neurons contributes to an excitatory descending pathway within the PAG–RVM–spinal dorsal horn neuraxis. By contrast, local GABAergic SST+ vlPAG neurons may participate in local circuit mechanisms related to defensive-state expression, including freezing. Further anatomical and functional studies will be required to resolve these possibilities.”

      In the absence of a direct identification of these cells exclusively releasing GABA, an alternative explanation should be considered. What about looking at vlPAG SOM+ neurons as a putatively mixed bag of local, inhibitory interneurons and long-range, RVM-projecting excitatory cells? This model would then open up interesting questions as to the actual function of somatostatin as a modulator of vlPAG circuit activity and associated function, and from my perspective, would nicely fit into the view of PAG circuits as integrators of complex survival responses.

      We thank the reviewer for this insightful suggestion and agree that, in the absence of direct evidence that vlPAG SOM+/SST+ neurons are exclusively GABAergic, an alternative interpretation should be considered. In particular, we agree that this population may be heterogeneous and could include both local inhibitory interneurons and long-range excitatory neurons projecting to the RVM. We believe this is an important and constructive framework for interpreting our data, and we have revised the Discussion accordingly. In the revised text, we now explicitly acknowledge the likely heterogeneity of vlPAG SST+ neurons and discuss the possibility that distinct local and long-range SST+ subpopulations may contribute differently to defensive-state regulation and descending pain modulation. We agree with the reviewer on this point and have modified the discussion accordingly (see point above).

      (2) "Our data indicate that the optogenetic inhibition of SST+ vlPAG cells promotes analgesia irrespective of the animal's defensive state. In contrast, the optogenetic activation of long-range SST+ vlPAG cells that project to the rostral ventromedial medulla (RVM) abolishes the analgesia mediated by fear behavior." (lines 32-35). Consider toning down these conclusions, as contrasting activation with inhibition of two different (though overlapping) populations cannot be fully conclusive. Alternatively, a pathway-specific (vlPAG-RVM) inhibitory experiment could help to fully understand the circuit mechanism and verify the necessity of these neurons.

      We thank the reviewer for raising this point. We agree that inhibition of the entire SST+ vlPAG population and activation of the long-range SST+ vlPAG neurons projecting to the RVM population are not directly equivalent manipulations. Our conclusion was intended at the level of observed functional effects: inhibition of SST+ vlPAG neurons promotes analgesia regardless of the defensive state, while activating long-range SST+ vlPAG neurons projecting to the RVM suppresses fear-conditioned analgesia. This occurs regardless of whether the SST vlPAG neurons are excitatory or inhibitory. To address the excitatory or inhibitory nature of SST vlPAG neurons, we have revised the discussion to include a reference to the Zhang et al study.

      (3) Despite an overall very thorough reporting style, some information is missing from the manuscript:

      (a) In Figures 2d and f, what are the freezing levels during optogenetic manipulation? From Figure 3d, one can expect that freezing is inhibited during the hot plate test, which could bias the NC response towards shorter latencies.

      We thank the reviewer for this important comment. As shown in Figure 1e, we previously quantified freezing both at CS onset and at the time of the nociceptive response in the hot plate test. These analyses indicate that freezing levels at the time of the nociceptive response do not differ between the CS+ and CS− conditions. Therefore, the variation in hot plate response latency is unlikely to be due to differences in freezing at the time of response.

      We acknowledge, however, that freezing was not directly measured during optogenetic manipulation in this experiment. Based on the temporal profile of freezing shown in Figure 1e, we still consider it unlikely that the effect of optogenetic manipulation on nociceptive latency is mainly caused by a change in freezing behavior.

      (b) In Figure 5, the histological experiment showing the vlPAG-to-RVM pathway is presented by a qualitative image only. Here, some quantification would strengthen the finding.

      We thank the reviewer for this comment. The aim of the histological experiment in Figure 5 was to provide qualitative anatomical evidence that vlPAG projections reach the RVM and are positioned in close apposition to spinally projecting RVM neurons. We did not intend this experiment to serve as a quantitative characterization of connectivity. We agree that a more systematic quantification would be informative, but this would require additional dedicated experiments beyond the scope of the present manuscript.

      (c) In Figures 6 c and d "Consistently, activation of the SST+ vlPAG-RVM pathway during CFCA had no impact on CS-presentation, whereas the same manipulation performed during CS+ blocked the increase in NC response latency compared to GFP controls." (line 194-196). Is it possible that the NC response cannot be any lower than the one during CS-, thus constituting a floor effect?

      We are thankful to the reviewer for this important point. We agree with the reviewer that this is indeed a possibility. We have added a sentence in the discussion to acknowledge this limitation.“Another possibility is that our nociceptive test with a slow ramp of temperature induces a floor effect on nociceptive response latency, which may limit the detection of further decreases in latency under certain conditions.”

      (c) Connected to major point 1- this experiment is important for defining the circuit mode and therefore should be as convincing as possible. However, for the colocalization experiment in Supplementary Figure 3, the methodological description is missing and thus makes it hard to comprehend how this data set was generated (how many data points, etc.). The visual depiction of the results is non-standard and not easily graspable. Consider e.g., a Venn diagram.

      We apologize for this omission in the original manuscript. We have now provided this methodological information in the method section. We have now expanded the description of these data in the figure legend to ease the comprehension of the figure.

      Reviewer #3 (Public review):

      Summary:

      Conditioned analgesia refers to the ability of a learned fear cue to suppress pain-related behavior and neural activity. Understudied, the authors developed a novel conditioned analgesia procedure in which a cue that had been paired or unpaired with shock was played while a hot plate increased temperature. Compared to several control conditions, the authors found increased latency to a nociceptive response (paw licking). The authors identified somatostatin neurons in the periaqueductal gray as a likely mediator of the behavior. They then showed that: (1) stimulating vlPAG-SST neurons blocked nociceptive response latency increases to the CS+, (2) stimulating vlPAG-SST neurons suppressed fear retrieval freezing, (3) stimulating vs. inhibiting vlPAG-SST neurons drove opposing modulation of c-fibers and Aδfibers, (4) direct-projecting vlPAG SST neurons modulate freezing while RVM-projecting vlPAG SST neurons modulate conditioned analgesia.

      Strengths:

      These experiments have many strengths. The behavioral assay is chief among them. The assay is robust and controls for confounding factors to reveal a repeatable effect of a shock-paired cue to delay nociceptive responding. The optogenetic experiments provide the correct level of temporal precision, given the authors' time-specific interest in cued responding. Combining neuronal manipulations with spinal recordings is particularly innovative, especially in the context of more behavioral neuroscience-based assays. All-in-all, I found this to be an exceptionally strong set of experiments.

      Weaknesses:

      No obvious weaknesses were identified by this Reviewer.

      Recommendations for the authors:

      Comments from Reviewing Editor:

      Summary

      Three reviewers have assessed your manuscript on vlPAG somatostatin pathways contributing to conditioned analgesia. Conditioned analgesia refers to the ability of a learned fear cue to suppress pain-related behavior and neural activity. Understudied, the authors developed a novel conditioned analgesia procedure in which a cue that had been paired or unpaired with shock was played while a hot plate increased temperature. Compared to several control conditions, the authors found increased latency to a nociceptive response (paw licking). The authors identified somatostatin neurons in the periaqueductal gray as a likely mediator of the behavior. They then showed that: (1) stimulating vlPAG-SST neurons blocked nociceptive response latency increases to the CS+, (2) stimulating vlPAG-SST neurons suppressed fear retrieval freezing, (3) stimulating vs. inhibiting vlPAG-SST neurons drove opposing modulation of c-fibers and Aδ-fibers, (4) direct-projecting vlPAG SST neurons modulate freezing while RVM-projecting vlPAG SST neurons modulate conditioned analgesia.

      Strengths

      All three reviewers converged on multiple strengths. The assay developed was seen to be novel, rigorous, and included a variety of controls that convincingly demonstrated conditioned analgesia. Focusing on the ventrolateral periaqueductal gray, and more specifically on somatostatin-expressing cells, made prior sense, and the results more than justified this selection. Approaching the vlPAG and circuits with many converging methods provided further, compelling evidence for a role in conditioned analgesia.

      Weaknesses

      Specific weaknesses are described in the individual reviews. Generally, the following weaknesses were identified. The study only used male mice, a choice that should be better justified. Animals were reasonably excluded from analysis, but the final group ns for analyses were not always clear. Some statistical results lacked clarity. The relevance of these findings to prior work (particularly Zhang et al. 2023, Journal of Pain) was not always described. Relatedly, the results would be better contextualized by appreciating and describing the likely diversity of somatostatin functional types and projection types.

      Recommendations

      (1) Provide rationale for only using male mice, discuss the limitation of the exclusion of females, and note that male mice were the subjects in the abstract.

      Thank you for this recommendation, we have mentioned this information in the abstract and in the discussion. We have also mentioned the limitations of not including female mice in the abstract and the discussion of the revised manuscript.

      (2) Complete final report ns for each statistical analysis. If you have not already done so, please include full statistical reporting including exact p-values wherever possible alongside the summary statistics (test statistic and df) and, where appropriate, 95% confidence intervals. These should be reported for all key questions and not only when the p-value is less than 0.05 in the main manuscript.

      An extended table with all statistical tests and analysis for all figures has been provided in sup Table 1.

      (3) Include example videos of CFCA sessions, demonstrating optogenetic effects.

      We understand the editor’s request to include video material illustrating the behavioral responses. However, we would prefer not to include such videos in the manuscript, in accordance with our institution's guidelines and recommendations on the dissemination of animal experimentation footage. Importantly, all behavioral sessions were systematically video-recorded from both sides of the apparatus, allowing detailed offline analysis of the animals’ responses. These recordings were carefully examined by an experienced experimenter to assess nociceptive behaviors, including jumping responses and licking of the stimulated hindpaw. This procedure ensured a reliable and accurate evaluation of pain-related behavioral reactivity. While the videos themselves cannot be included in the manuscript for the reasons mentioned above, we believe that the behavioral scoring procedures described in the Methods section provide a clear and rigorous description of how these responses were assessed. In addition, Figure 1 includes an example image illustrating hindpaw licking behaviour, which is typically more subtle and more difficult to identify than jumping responses. We therefore believe that this visual example, together with the detailed description of the scoring procedure and the quantitative data provided, adequately supports the interpretation of the behavioural results.

      (4) Provide summary expression and ferrule placement figures.

      We thank the editor for this comment. We have now included schematic summaries of fiber placements for both SST and VIP mice used in this study, based on histological verification (Supplementary Figures 10 and 11). Representative images of viral expression are also provided (Figure 2a, Supplementary Figure 7b and f).

      (5) Detail how behavior judgments were made.

      We thank the editor for emphasizing this important methodological point. During all behavioral sessions, mice were video-recorded simultaneously from both sides of the apparatus, allowing a comprehensive and unobstructed view of the animals’ posture and movements throughout the experiment. These recordings were subsequently analyzed offline by an experienced experimenter trained to evaluate nociceptive behaviors. Pain-related behavioral responses were assessed based on well-established indicators of nociceptive reactivity. In particular, we quantified overt escape-like reactions such as jumping, which reflects a strong aversive response to the stimulus. In addition, we evaluated more localized nociceptive behaviors directed toward the stimulated limb, including licking of the hindpaw. These measures are commonly used in rodent pain assays and provide reliable behavioral readouts of nociceptive sensitivity. The combination of bilateral video recordings and expert behavioral scoring ensured that both subtle and robust nociceptive responses could be accurately detected and categorized during the analysis.

      (6) Provide the temperature at which nociceptive responses were initiated. Check grammar and references.

      The temperature at which nociceptive responses were initiated were originally reported in Supplementary Figure 1, 2 and 5.

      Reviewer #1 (Recommendations for the authors):

      (1) The authors use optogenetic manipulation of SST activity in the vlPAG to show that this cell type is involved in fear-induced analgesia. They include a valuable control to show that manipulation of another inhibitory cell type (VIP) also does not impact analgesia. It would be helpful to know the expression level of VIP cells in the vlPAG. Is this a predominant inhibitory projection cell in the vlPAG (besides SST)?

      We thank the reviewer for pointing this. While we did not quantify the expression level of VIP+ cells in the vlPAG in the present study, available data suggest that this population is relatively sparse compared to other inhibitory cell types. In particular, reference to the Allen brain atlas indicates that VIP gene expression in the vlPAG is limited and primarily localized around the fourth ventricle, within the lateral and ventrolateral PAG, rather than broadly distributed across the region. Consistent with this, we provide an example of viral expression in VIP-Cre mice in Supplementary Figure 7f, illustrating the restricted distribution of VIP+ neurons in the vlPAG. We have also provided a summary of ferrules placement for SST and VIP mice used in our study in Supplementary Figures 11 and 10, respectively.

      (2) The numbers of animals dropped from each experiment should be indicated - perhaps on the statistics table?

      We thank the reviewer for pointing this.

      As stated in the Methods, we applied strict inclusion criteria for mice undergoing the hot-plate test, specifically a discrimination index ≥ 0.4 and a conditioning index ≥ 0.3. Using these criteria, 23% of wild-type mice were excluded for failing to meet the discrimination criterion. In the transgenic groups, an average of 20% of mice failed to meet the learning criteria, and an additional 12% were excluded due to incorrect opsin injection or misplaced optic fiber placement.

      (3) Line 105: "...,which activity..." change to "..., whose activity..."

      Done

      Reviewer #2 (Recommendations for the authors):

      (1) Please also provide absolute temperature values of the nociceptive response threshold.

      The temperature at which nociceptive responses were initiated was originally reported in Supplementary Figure 1, 2 and 5.

      (2) It would be nice to see an example video of a CFCA session (with and without optogenetic manipulation).

      We understand the editor’s and reviewer’s request to include video material illustrating the behavioral responses. However, we would prefer not to include such videos in the manuscript, in accordance with our institution's guidelines and recommendations on the dissemination of animal experimentation footage. Importantly, all behavioral sessions were systematically video-recorded from both sides of the apparatus, allowing detailed offline analysis of the animals’ responses. These recordings were carefully examined by an experienced experimenter to assess nociceptive behaviors, including jumping responses and licking of the stimulated hindpaw. This procedure ensured a reliable and accurate evaluation of pain-related behavioral reactivity. While the videos themselves cannot be included in the manuscript for the reasons mentioned above, we believe that the behavioral scoring procedures described in the Methods section provide a clear and rigorous description of how these responses were assessed. In addition, Figure 1 includes an example image illustrating hindpaw licking behaviour, which is typically more subtle and more difficult to identify than jumping responses. We therefore believe that this visual example, together with the detailed description of the scoring procedure and the quantitative data provided, adequately supports the interpretation of the behavioural results.

      (3) Please provide a schematic summary of fiber placements and opsin expressions confirmed by histological examinations.

      We thank the reviewer for this comment. We have now included schematic summaries of fiber placements for both SST and VIP mice used in this study, based on histological verification (Supplementary Figures 10 and 11). Representative images of viral expression are also provided (Figure 2a, Supplementary Figure 7b and f).

      (4) "Valid nociception readout responses included jumping or licking the hindpaw." (Line 453). How was this evaluated- manually or automated, blinded etc.?

      We thank the reviewer for emphasizing this important methodological point. During all behavioral sessions, mice were video-recorded simultaneously from both sides of the apparatus, allowing a comprehensive and unobstructed view of the animals’ posture and movements throughout the experiment. These recordings were subsequently analyzed offline by an experienced experimenter trained to evaluate nociceptive behaviors. Pain-related behavioral responses were assessed based on well-established indicators of nociceptive reactivity. In particular, we quantified overt escape-like reactions such as jumping, which reflects a strong aversive response to the stimulus. In addition, we evaluated more localized nocifensive behaviors directed toward the stimulated limb, including licking of the hindpaw. These measures are commonly used in rodent pain assays and provide reliable behavioral readouts of nociceptive sensitivity.The combination of bilateral video recordings and expert behavioral scoring ensured that both subtle and robust nociceptive responses could be accurately detected and categorized during the analysis.

      (5) Line 226 REF33 doesn't seem to fit.

      The reference list has been updated. Related to this section in which we discuss the disinhibition mechanisms inducing nociception in chronic stress mice. We have cited the work of Samineni et al., 2015 (reference 15) and Tovote el al., (reference 23) both related to these disinhibition mechanisms.

      Full sentence for reference 33 (now 35): “Two independent previous studies found that long-range inhibitory inputs from the central medial amygdala contact inhibitory cells within the vlPAG, implicated in different roles: the modulation of fear behavior (23) and nociceptive transmission (35)”.

      Ref 35 - Yin, W. et al. A Central Amygdala–Ventrolateral Periaqueductal Gray Matter Pathway for Pain in a Mouse Model of Depression-like Behavior. Anesthesiology 132,1175–119 (2020)

      (6) Some minor language, semantic, and grammatical flaws.

      The manuscript has been evaluated for language, semantic and grammatical flaws

    1. eLife Assessment

      This important work challenges current models of merozoite surface protein function by showing that MSP2 is dispensable for parasite growth while modulating immune responses to AMA1, with implications for malaria vaccine design. The conclusions are supported by compelling experimental evidence, including state-of-the-art technologies and well-characterized monoclonal antibodies. These findings provide new insights into immune evasion and antigen targeting that will be of broad interest to parasitology, immunology, and vaccine researchers.

    2. Joint Public Review:

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      The major strengths of the manuscript are in the Plasmodium falciparum genetic and phenotyping approaches. PfMSP2 knockouts are made in two different strains, which is important as it is know that invasion pathways can vary between strains, but is a level of comprehensiveness that is not always delivered in P. falciparum genetic studies. The knockout strains are characterised very thoroughly using multiple different assays and the authors should be commended for publishing a good deal of negative data, where no phenotype was detected. This is not always done but is very helpful for the field and reduces the potential for experimental redundancy, i.e., others repeating work that has already been performed but never published. The quality of the writing, referencing and figures is also generally strong.

      There are certainly some areas of the manuscript that would benefit from deeper exploration, such as electron microscopy/other imaging approaches to explore whether deletion of PfMSP2 has a visible impact on merozoite surface structure, further replicates of the video microscopy assays to see whether trends in the data could reach significance (although these are very time-consuming and technically difficult assays), and follow up of some of the genes where expression is changed by PfMSP2 knockout (as the authors point out, there are no candidates that have a very obvious link to invasion suggesting that they may be compensating for PfMSP2 function, although several are expressed in schizont stages). However, there is already a substantial amount of data in the manuscript, and more detailed follow-up is reasonable to leave to future work. Overall, with the modifications made through the review process, including the addition of new controls for key experiments, the claims and conclusions are justified by the data, and the manuscript generates important new information about a highly studied Plasmodium falciparum merozoite surface protein. The studies are important and have potential for directing vaccine design targeting erythrocyte invasion, a critical step in bloodstream expansion of malaria parasites.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      (1) There are certainly some areas of the manuscript that would benefit from deeper exploration, such as electron microscopy/other imaging approaches to explore whether deletion of PfMSP2 has a visible impact on merozoite surface structure.

      We in principle agree with the reviewer that applying enhanced resolution microscopy approaches to understand structural and functional changes with loss of PfMSP2 could be of interest. However, based on our ongoing work, this represents a significant body of work in terms of experimental optimisation in an effort to gain the detail required to make meaningful insights. Therefore, this will remain outside the scope of this manuscript and we hope to provide these insights in future studies.

      (2) Further replicates of the video microscopy assays to see whether trends in the data could reach significance (although these are very time-consuming and technically difficult assays).

      Conclusions we have drawn from live-cell imaging data for MSP2 knock-out parasites encompass some 43 invading merozoites from 21 schizont ruptures for PfDd2 WT and 35 invading merozoites from 18 schizont ruptures for PfDd2 DMSP2 parasites. One of the leading studies to apply live-cell microscopy to film invading merozoites based conclusions of invasion kinetics on: 3D7 (number of merozoite invasion =63, number of schizont ruptures =23), D10 (invasions =33, ruptures =20) and W2mef (invasions =39, ruptures = 15; this line is of the same lineage as Dd2) (Weiss et al. PLoS Pathogens, 2015). Although there are variations within and between lines from this gold-standard study, our dataset is mostly comparable in terms of the number of schizont ruptures and merozoite invasions filmed and analysed to look at changes in kinetics. What we can say definitively is that there is no strong phenotype in the absence of inhibitory antibodies against other antigens for either live-cell or growth inhibition assays. Therefore, we have focussed the data interpretation in the manuscript to highlight the lack of statistical significance and limited phenotype seen, which given the previously believed importance of MSP2 to P. falciparum invasion of red blood cells is somewhat surprising.

      In order to address this suggestion, we have modified the discussion to better represent any non-significant changes in invasion and growth seen.

      “Despite the abundance of PfMSP2 on the merozoite surface and previous work suggesting a role in RBC invasion, we found merozoites invade and grow with similar kinetics to wildtype parasites in the absence of PfMSP2. This does not exclude a role for PfMSP2 in vivo where there are additional pressures, such as immune-effector mechanisms and flow dynamics, on merozoite invasion. However, given we have knocked-out PfMSP2 from two different P. falciparum isolates, our findings do not currently support a major role for PfMSP2 in the mechanics of merozoite invasion. Thus, it appears that the function of the two most abundant proteins on the merozoite surface, PfMSP1 (Das et al., 2015; Kals et al., 2024) and PfMSP2, are not obviously linked to merozoite binding to the RBC and subsequent invasion.”

      (3) Follow up of some of the genes where expression is changed by PfMSP2 knockout (as the authors point out, there are no candidates that have a very obvious link to invasion suggesting that they may be compensating for PfMSP2 function, although several are expressed in schizont stages).

      A thorough investigation of the genes where expression changes with PfMSP2 knock-out would require a substantial body of additional work, not least because they would all have to be investigated as there is no single likely candidate based on stage of expression, membrane binding properties or previous links to merozoite surface architecture. Given this, potential follow up of these proteins will be left for future studies.

      We also thank the reviewer for the recognition of the work provided in the manuscript and the modifications made that have improved the manuscript from version 1. The reviewer also recognises the value in our detailed characterisation, including data where phenotyping changes with MSP2 knock-out could not be seen, in defining the function of PfMSP2 as commented below:

      However, there is already a substantial amount of data in the manuscript, and more detailed follow-up is reasonable to leave to future work. Overall, with the modifications made through the review process, including the addition of new controls for key experiments, the claims and conclusions are justified by the data, and the manuscript generates important new information about a highly studied Plasmodium falciparum merozoite surface protein.

      Reviewer #3 (Public review):

      Major points:

      (1) Much of the manuscript describes negative results and this reviewer found it arduous to get through many negative or nonsignificant results before finally getting to the significant effect on AMA1 inhibitory antibodies, not presented until Figure 6! Computational studies in Fig. 1 could be a supplementary figure. Figs. 2 and 3. demonstrate knockout in 3D7 and Dd2, respectively and could be assembled into a single figure. (Notably Fig. 2A and 3A are almost identical with use of some different primers.) Fig. 2E, 2F, 3D-H, all of Fig. 4, most of Fig. 5 are all negative or insignificant results that could also be moved to supplementary data. As MSP4, MSP5, and SUB1 are presumably included in the whole genome RNA-seq experiments shown in Fig. 4C, it makes sense to remove Fig. 4A data from the paper fully. These consolidating changes would help highlight the key finding of improved binding and block of AMA1's role in invasion.

      We have chosen to not take the approach proposed by Reviewer 3 as it would leave the manuscript with only around 2.5 Figure panels and undersells the very significant amount of work that has been done to characterise PfMSP2 knock-out lines. Although, as noted by the reviewer, piggyBac mutagenesis studies predict PfMSP2 is dispensable, much of the field likely expect PfMSP2 to be essential to P. falciparum blood stage parasite growth due to the results of earlier reverse genetics approaches and many years of publications that have speculated on the importance of the protein. Therefore, we are also conscious of providing very clear and comprehensive evidence to support our findings. While this may delay highlighting the findings in Figure 6, we also note that the lengths we have gone to in characterising an important antigen with a difficult phenotype is still valued as evidenced by Reviewer 2 (Public Review Comments on the original manuscript):

      “PfMSP2 knockouts are made in two different strains, which is important as it is known that invasion pathways can vary between strains, but is a level of comprehensiveness that is not always delivered in P. falciparum genetic studies. The knockout strains are characterised very thoroughly using multiple different assays, and the authors should be commended for publishing a good deal of negative data, where no phenotype was detected.”

      (2) The potentiating effects on anti-AMA1 antibodies are shown with rabbit sera and purified antibodies, mouse monoclonal antibodies, and smaller i-bodies inspired by shark antibody-like receptors but not with human monoclonal antibodies (hmAbs). As naturally acquired hmAbs targeting AMA1 have been identified and characterized (PMIDs: 39632799, 40020675), would it not be important to test these antibodies in the ∆MSP2, especially as the authors emphasize the importance of their model in designing better human malaria vaccines?

      As the reviewer noted, we demonstrated enhanced inhibitory activities of antibodies to AMA1 using rabbit polyclonal antibodies, mouse mAbs, and i-bodies. We note that the WD34 i-Body we used was humanised to be IgG-like with a human Fc-region (IgG1 backbone). Rabbit IgG is very similar to human IgG1. Therefore, we have provided evidence of the enhancing effect using different types and sources of antibodies relevant to human immunity to support our conclusions. Our findings open new avenues for future research and we agree with the reviewer that future studies using panels of human mAbs to defined epitopes would be interesting and may further inform vaccine design; however this is beyond the scope of the current paper. We do not have the mAb mentioned by the reviewer to test in our system. To perform studies with human mAbs would take a substantial amount of time (many months), requiring the generation of different human mAbs and quantification of their activity and testing them for potentiation effects. While this would be an interesting future endeavour, we do not feel that such studies are needed at this stage to support our conclusions, and instead would be a future extension from our current paper. To acknowledge the reviewer's comment, we have extended our comment in the discussion about future studies with different panels of invasion inhibitory antibodies to include huMabs targeting AMA1 as follows:

      “Further investigation using the parasite lines developed in this study and a wider panel of antibodies that target different stages of the merozoite invasion process, including human monoclonal antibodies against AMA1 (Patel et al., 2025), could shed more light on this potentially novel mechanism of vaccine derived antibody efficacy.”

      (3) Fig. 7 presents quantitative fluorescence microscopy to measure anti-AMA1 binding and support a model where MSP2 serves to sterically hinder antibody access to AMA1 on individual merozoites. I understand that the negative WD33 control is useful to contrast to the positive WD34 antibody (both bind AMA1 but only WD34 exhibits parasite growth inhibitory effects), but it seems that use of smaller i-bodies rather than conventional larger mouse or ideally human monoclonal antibodies may compromise demonstration of steric hindrance by MSP2 because smaller i-bodies may be less hinder.

      The antibodies used in this experiment have fluorescent tags attached. So while the untagged WD33 and WD34 i-bodies are approximately 14 kDa, when fused to GFP or mCherry their expected size increases to approximately 42 kDa, approaching that of the Fc-tagged WD34 i-body (78 kDa) that shows increased growth inhibitory activity in the absence of MSP2. Therefore, we expect steric hindrance to be a significant factor with these fluorescently tagged antibodies.

      (4) Some explanation for why WD33 fails to inhibit growth despite targeting the same antigen as WD34 is needed. Are the epitopes known? Does one bind further from the RON2 binding pocket?

      As reported in Angage et al., Nature Communications 15, 7206 (2024). WD34 has been identified to bind to, and block, a site within the hydrophobic AMA1 and RON2 binding pocket found on Domain II of AMA1. In contrast, WD33 recognises a distinct conserved epitope in Domain II of AMA1 near to, but not overlapping with, the hydrophobic AMA1 and RON2 binding pocket. We have clarified this by including additional description when first describing the i-bodies as follows:

      “When we tested the i-body WD34 (Angage et al., 2024) which binds a highly conserved epitope that includes the PfRON2-binding pocket on PfAMA1 domain II, we observed a small potentiation of PfAMA1 specific activity with knock-out of PfMSP2 in Pf3D7 (1.3-fold; IC<sub>50</sub> PfD7 WT 0.012 mg/mL; IC<sub>50</sub> Pf3D7 DMSP2 0.009 mg/mL; p=0.08 Figure 6F).”

      Then

      “A second i-body, WD33 (Angage et al., 2024), which binds AMA1 between domain II and domain III but does not appear to overlap with the PfRON2-binding pocket on PfAMA1, had very limited invasion inhibitory activity against Pf3D7 parasites and did not show improved potency with knock-out of Pf3D7 MSP2 (0.9-fold; IC<sub>50</sub> Pf3D7 WT 1.02 mg/mL; IC<sub>50</sub> Pf3D7 DMSP2 1.1 mg/mL; p=0.8; Figure 6I).”

      Recommendations for the authors:

      Reviewing Editor Recommendations:

      Although providing microscopic images might require a lengthy process, including results based on human mAbs (if available) might enhance the strength of evidence. The reorganization of the figures and the presentation of results usually falls into the realm of personal preferences, however, if the comments/suggestions are useful, it might highlight your message.

      As covered in the Response to Public Reviewer Comments for Reviewer 2 and indicated by the editor, investigations of phenotypes found in this study using high-resolution imaging techniques (e.g. electron microscopy) will require very significant additional work and will be attempted in future studies. We also provide a response to Reviewer 3 in regards to the potential to test human monoclonal antibodies and believe this is best done more thoroughly in future studies. We have elected to not make substantial changes to the data presented as suggested by Reviewer 3. We have addressed additional comments as covered below.

      Reviewer #3 (Recommendations for the authors):

      Minor Comments

      (1) Scale bar in Fig. 7A is not resolved well. The image is too pixelated to resolve merozoites or the actual dimensions of the scale bar.

      We have updated this figure to provide improved clarity of the scale bar.

      (2) Lines 69, 216, 221, 253, 628-629, 648 all suggest that MSP2 was heretofore assumed to be essential. However, piggyBac insertional mutagenesis revealed that MSP2 is highly dispensable (MIS of 0.988, per PlasmoDb.org; PMID: 29724925). I would suggest to tone down this claim as it does not detract from the authors' production of useful ∆MSP2 clones.

      We agree with the reviewer that the piggyBac insertional mutagenesis study results should also be acknowledged and apologise for this oversight. To address this, we have reviewed the sentences highlighted by the reviewer and, where appropriate for the historical interpretation of PfMSP2 function, have added the following modified information through the text:

      P. falciparum merozoite surface protein 2 (PfMSP2), an antigen reported to be refractory to gene knock-out in P. falciparum (Sanders et al., 2006) but that has also been reported to be dispensable in a piggyBac mutagenesis study (Zhang et al., 2018), has been of long-term interest as a vaccine candidate.”

      “Given previous unsuccessful attempts to disrupt pfmsp2 (Sanders et al., 2006), and its high abundance on the merozoite surface (Gilson et al., 2006), PfMSP2 has been traditionally viewed as an essential P. falciparum protein with an essential function in merozoite invasion, although more recent piggyBac mutagenesis studies have called this understanding into question (Zhang et al., 2018).”

      We have chosen not to modify this text and it remains the same as below. The reason for not changing this text is the result that we could knock-out MSP2 from 3D7 was still unexpected given the published reverse genetics studies and results from piggyBac mutagenesis studies are also sometimes not reliable indicators of what happens when reverse genetics is performed. Therefore, the following text we believe is a reasonable description.

      “Unexpectedly, we confirmed successful disruption of pfmsp2 by replacing the coding sequence between 132 bp and 819 bp of the gene with a hDHFR drug selection cassette in the 3D7 P. falciparum laboratory-adapted line (Figure 2A and B), resulting in Pf3D7 DMSP2 parasites.”

      “As a previous reverse genetics study in 3D7 reported that PfMSP2 was essential for P. falciparum growth in vitro (Sanders et al., 2006), we investigated whether PfMSP2 could also be removed from PfDd2, an isolate of P. falciparum that differs from 3D7 in geographical origin, RBC receptor usage and allelic type of pfmsp2.”

      “However, CRISPR-Cas9 gene editing used in this work has shown that, in contrast to previous attempts to knock-out PfMSP2 (Sanders et al., 2006), PfMSP2 is not essential for P. falciparum blood stage parasite growth in vitro.”

      “Advancements in gene-editing techniques in P. falciparum have allowed us to directly demonstrate using reverse genetics in two different parasite lines that PfMSP2 is not essential for P. falciparum growth in vitro.”

      (3) Figs. 2B, 2C, 2D show PCR, immunoblots, and IFA with a ∆MSP2 clone but two clones (termed clone 1 and clone 2) are show in panels 2E and 2F. Which clone is used in each panel? Without clarification, readers may wonder if one clone was used for PCR but another clone gave a desired result in immunoblots? By convention, validation studies (PCR and immunoblots) should be performed and shown (in Supplementary figures) for all clones used for phenotype studies; alternatively, a single clone can be used throughout if all clones are presumed identical. Which of these clones was used for the RNA-seq experiments in Fig. 4C? Similar questions arise for the two knockout clones made in the Dd2 line (Fig. 3D).

      We agree with the reviewer that it would be helpful to have this information provided more clearly through the Results. To this end, we have updated the Figure legends across Figures 2, 3, 4, 5, 6, 7 and Supplementary Figure 5 as appropriate to specifically indicate the clones used for the downstream experiments. All clones were validated by PCR and, after growth characteristics were found to be the same, a single clone was used for all downstream experiments for PfMSP2 knock-outs in both 3D7 and Dd2.

    1. eLife Assessment

      This important work addresses a very relevant biological question: what is the cellular basis of wound healing? Using the Drosophila pupal notum as a model, the paper provides an elegant, thorough, descriptive characterization of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors meticulously characterize the cell-cell fusion events during wound healing and inhibit cell fusion to show to that it is necessary to speed wound closure. In addition, the study provides convincing evidence that cell fusion allows actin resources at be partitioned to the leading edge.

    2. Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental yet poorly understood biological process.

      Comments on revised version.

      The manuscript overall is significantly improved and authors addressed majority of my concerns. The addition of the computational vertex model (Figure 7) as well as Atg1 RNAi (Figure 4) to inhibit cell fusion provide more mechanistic insight to their study. However, the analysis of Atg1 RNAi wound assay falls short as it does directly measure changes in syncytium frequency nor size to confirm that cell fusion is reduced. The authors should quantify the number of nuclei per syncytium over the 2hr wound healing period as performed for WT in Figure 1C. It would have been ideal if they could have also performed the Act-GFP spreading assay in WT and Atg1 RNAi strains to determine if Act-GFP movement is dependent on cell fusion as purposed. At the least, further quantification of Atg1 RNAi phenotype is warranted to support their conclusions.

    3. Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laser-induced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFP-positive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to syncytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Comments on revised version:

      The authors have extended their original manuscript by adding two key parts. First, they show a role of Atg1 in mediating cell fusion (Figure 4). Second, they provide additional evidence for a contribution of radial border fusions to wound closure through its effect on tissue fluidity and through computational modelling (Figure 7).

      This new version of the manuscript is greatly improved and provides significant new insights into the role of syncytia in aiding wound repair. There are just a few minor, yet important, additions needed to back up Figure 4 which should not require new experiments.

      Minor but important points:

      The authors show a role of Atg1 in mediating syncytia formation in Figure 4. However, since the Pnr>+ side of the wound closes slower than the non-Pnr side (control side), a few additions to this figure would be important and should not require additional experiments.

      (1) The authors should show, similar to the data shown in Figure 4D of the wound radius over time for control versus Pnr>Atg1RNAi, also the same type of data for control versus Pnr>+.

      (2) Since Pnr>+ also slows down wound healing, albeit to a lesser extent than Pnr>Atg1, the authors should also show an extra graph that provides evidence that Pnr>Atg1RNAi reduces syncytia formation more than Pnr>+ does. E.g. Two graphs could be added that show individual cell size at 4 or 5h post wounding for control versus Pnr>Atg1RNAi as well as for control versus Pnr>+ and also another graph with the same data but comparing cell size between Pnr>+ and Pnr>Atg1RNAi. Otherwise, if the expected minimum cell size for a syncytium is easy to estimate, a graph could be added that shows the percentage of cells that are above this threshold (e.g. above 100 square micron) for control versus Pnr>Atg1RNAi and control versus Pnr>+ and Pnr>+ versus Pnr>Atg1RNAi.

    4. Reviewer #3 (Public Review):

      In this revised manuscript, White et al. aimed to understand the wound-induced syncytia formation behavior in wound repair of Drosophila melanogaster pupal notum. For this purpose, the authors characterized two different types of adherens junctions' outcomes during syncytia formation around the wound region - border breakdown versus apical shrinking which appear to happen in different time points and for different time durations. The authors characterized cell-cell fusion events using cytoplasmic, junctional and nuclear markers. They determined that about half of the cells within 70 um radii from the wound undergo cell-cell fusion. They studied wound induction on the border between control epithelia and pnr domain suggesting that Atg1 is required for post-wound syncytia formation and wound closure. They showed that during wound closure syncytia gradually invade the wound leading edge mostly by radial fusion events. The data suggests that intercalation of cells from the leading edge slows down the wound closure process. They propose that cell fluidity of syncytial cells plays a role in wound closure speed. Finally, the authors showed that actin is concentrated to the front edge of syncytia located in the wound leading edge. The authors described some aspects of syncytia formation during wound closure using different approaches. Some clarifications are needed as described below.

      Major suggestions:

      (1) Introduction, page 4. The examples of developmental syncytia formation of invertebrates and vertebrates are confusing. The authors may want to make the examples clear and add additional examples. Currently, readers may assume that C. elegans cell fusions occur only in the hypodermis - other structures can be mentioned like the vulva, pharyngeal muscles, glia, tail. In addition, the authors may want to add injury-induced fusions like the C. elegans' PLM and PVD neurons (Ghosh-Roy et al., 2010; Newman et al., 2015; Oren-Suissa et al., 2017).

      (2) In cases where it is not clear whether fusion has occurred or whether mononucleated cells were ejected from the leading edge, membrane markers can be used. Page 6. Lines 96-99. The authors may want to use a membrane marker like RFP-PH driven by the epithelial cell promoter.

      (3) Pages 8-10. The authors may want to clearly explain that apical junctions shrinking is a post fusion event. That the apical shrinking is caused by the expansion of fusion pores and the migration of apical junctions towards the basolateral domain. This is something that was clearly shown during physiological epidermal cell-cell fusion in C. elegans by Mohler et al., 1998 and 2002. A cartoon showing the process of cell-cell fusion, pore expansion and apical junction dynamics would make the manuscript much clearer.

      (4) Page 9. Line 170. "...as these cells represent fusion initiation events (fusion pore) but were unable to productively stabilize and expand the site of fusion and so returned to the diploid state." The authors may want to make clear that this is an assumption that needs to be tested. Live imaging using a membrane marker may resolve whether a reversible fusion pore was generated.

      (5) Page 11. It is not clear whether Atg1 is directly required for cell fusion, or that autophagy is required for efficient cell fusion or both Atg1 and autophagy participate in the fusion process.

      (6) Page 12. Line 235. "Indeed, we observed that several hours after wounding, the entire leading edge was occupied by syncytia." This observation is based only on the adherens junction marker. Can they test basal cell membrane marker? Is it possible that the mononucleate cell in the leading edge is under the two syncytia?

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental, yet poorly understood biological process.

      Weaknesses:

      A major weakness is that all the authors' conclusions are based on descriptive studies, in which the role of cell fusion is not directly tested. This is particularly important because other models of wound induced polyploidization have demonstrated that another cytoskeletal protein, myosin, was upregulated and dependent on endoreplication, and not cell fusion. Therefore it remains unclear to what extent cell fusion, endoreplication, or both are required to outcompete mononucleated cells as well as pool actin as described in this study.

      We thank the reviewer for appreciating our live imaging and meticulous approach. In this revision we have identified that the gene Atg1 is required for wound-induced fusion in the pupal notum: when Atg1 is knocked down, there is a reduction in wound-induced cell fusions, both border breakdown and cell shrinking. Analysis of Atg1 knockdown shows that the wounds close more slowly. This is a direct test of the role of cell fusion in speeding wound closure, presented in new Fig. 4.

      Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laserinduced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFPpositive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to sycytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Weaknesses:

      The authors provide some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound. The authors suggest that the syncytial cells might be better able to close the wound. However, some genetic studies would need to be done to establish this more convincingly. E.g. Could the authors genetically block syncytia formation and then show that these wounds now heal slower?

      We now present such data in new Fig. 4, which describes knocking down Atg1, previously shown by the Leptin lab to promote wound-induced fusions in larval epidermis. We quantify the resulting reduction in fusion in the pupal notum and show that the leading edge advances more slowly to heal the wound.

      The authors suggest that radial border breakdown reduces the requirement for cell intercalation. While this might be true it also raises the question of how the various syncytia facing the wound border change shape to allow the shrinkage of the first cell row over time to allow wound closure. None of the four movies included in the study shows the whole wound healing process until the later stages, making it hard to assess this. It would be good to include one such movie showing the syncytia in the whole wound and comment on this point.

      In response to the reviewer's request, we now extend Supplemental Video S1 out through 8 hours after wounding (same video as included previously but extended longer). In this video, as in many of the wounds, it is hard to determine the exact moment of closure because a syncytium extends across the wound whereas the nuclei do not. However, during the process of closure, one can clearly observe the large syncytia becoming more wedge-shaped – drastically reducing the section of their perimeter remaining in contact with the wound’s leading edge.

      In addition, we now explore how syncytia reduce the need for intercalation in a computational model, presented in new Fig. 7 and Supplemental Videos S5 and S6. One can observe the modeled syncytia becoming similarly wedge-shaped. The modeling shows that the presence of syncytia and their ability to reshape can speed closure by about 1/3 even if the syncytia have no special properties aside from their relative size.

      In both the experiments and models, some syncytia are also removed from the leading edge by intercalation, but the presence of syncytia reduces the total number of intercalations needed.

      The authors hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. They show convincingly through the fusion of a single Actin-GFP-positive cell in the second cell row with a GFP-negative cell in the first cell row that Actin-GFP spreads in the fused cell and labels the previously unlabelled actomyosin cable. While the hypothesis of resource sharing to improve healing is intriguing and makes sense, this experiment doesn't necessarily prove the benefit of resource sharing. It does show cytoplasmic mixing following fusion, now allowing the GFPlabelled actin to diffuse and be incorporated into the actomyosin cable. In a wild-type condition, fusion would not increase the total concentration of resources, although it would increase the total amount of resources within this bigger fused cell. The question is whether resource sharing without increasing the protein concentration is beneficial and increases the efficiency of certain wound healing mechanisms. There might be a benefit of cell fusion, if for example certain resources were only present in limited amounts or if protein transport could increase the concentration locally. To provide better evidence for the hypothesis that resource sharing improves wound healing, maybe the authors could look at the actomyosin cable in a wounded epithelium (such as in Figure 4E, F), in which all cells express MyoII-GFP. The authors could compare the average intensity of the actomyosin cable at the wound edge in mononucleated cells versus in syncytia. If resource sharing is indeed beneficial, it might be that the actomyosin cable is stronger/brighter in syncytia or it forms quicker.

      We agree with the reviewer that we have not "proved the benefit of resource sharing". Because we cannot inhibit resource sharing while still allowing cell fusion, we can think of no rigorous way to test this hypothesis. We appreciate the reviewer's suggestion of quantifying the myosin at the leading edge cable, but we can imagine too many caveats to the interpretation to make it worthwhile. Rather, we accept the limitation that this is an untested, perhaps untestable, hypothesis -- but nevertheless intriguing.

      We do want to clarify ideas about the concentration of resources after fusion. We agree that the overall concentration of a given resource (mass/volume) throughout a syncytium would be the same as the overall concentration in the unfused progenitor cells; however, a syncytium would have a larger total resource mass to direct subcellularly, allowing for local subcellular concentration to be greater in a syncytium vs. an unfused cell. We demonstrate this subcellular localization of actin in a syncytium twice, in Fig. 7C and E (previously Fig. 6C,E), which we think is evidence for increased local concentration.

      The biggest limitation of this study is that the authors don't address how the formation of these syncytia is regulated. While the manuscript in its current form provides some valuable new insights into syncytial-driven wound closure, it would be much more informative if it also provided some mechanistic details. The authors could test if some of the mechanisms shown to regulate syncytial formation in other types of syncytia-driven wound healing are also involved here. E.g. Yorkie was shown to negatively regulate cell fusion in adult syncytial-driven wound closure (Losick et al 2013). The authors could test for the effect of Yorkie-RNAi in the epithelium on wound closure and syncytia formation. Expression of the dominant negative RacN17 also blocked cell fusion in adult syncytial-driven wound closure (Losick et al 2013).

      Moreover, JNK activation was shown to be needed in larval syncytial-driven wound closure (Galko and Krasnow 2004). The authors could test JNK pathway reporters to assess pathway activation or test if the JNK pathway is needed for syncytial-driven wound closure by expressing a dominantnegative form of Basket JNK in the epithelium.

      Or could syncytia formation be regulated by changes in Integrin-mediated adhesion as shown by the Galko lab in Wang et al 2015? They show that wounding provoked a striking relocalization of PINCH and ILK, indicating the disassembly of functional FA complexes concomitant with syncytium formation. Maybe the authors could investigate some of these.

      We investigated the role of JNK in fusion by expressing bsk<sup>DN</sup> on one side of the wound. Comparing the numbers of border-loss fusion on each side, we did not find a significant difference in our seven-sample cohort (see Author response image 1). If we had increased the sample size, we may have found a significant difference with a small effect size, but because of the small difference in fusions on each side we did not think this was worth pursuing. Instead, we include data that the autophagy gene Atg1 is required for cell fusion in new Fig. 4, which begins to address mechanism, and relates the wound-induced fusion described here in pupae to wound-induced fusion shown in larvae. A complete mechanism for wound-induced fusion is outside the scope of this paper, as we focus on the function of syncytia in healing wounds.

      Author response image 1.

      Another general question that the authors raise but don't address enough is whether syncytia-driven wound closure in proliferation-competent epithelia is any different from the one in post-mitotic, polyploid epithelia. Since the mechanism regulating the former is not known, this remains unclear.

      We now include a paragraph on this question in the discussion.

      Finally, it is not clear, whether syncytia in these proliferation-competent epithelia get resolved after wound healing. Do they get removed and replaced by mononucleated proliferation-competent cells or do the syncytia stay in the epithelium like a scar? The authors should provide some images of wound areas a few hours after wound closure is complete and comment on this.

      To answer the reviewer’s question: some but not all syncytia do get removed during wound closure by remarkable apoptotic/extrusion events. This will be the subject of a future manuscript, as it is outside the scope of this paper focusing on the function of syncytia in promoting wound healing.

      Minor points:

      Figure 3: It would be better to have the microcopy images alongside the quantifications.

      The images in Figs. 1 and 2 show the border breakdown and shrinking cells, and we do not see benefit in adding them in Fig. 3.

      Figure 4A: The syncytium at the wound edge here doesn't look straight but wavy. Does it not form an actomyosin cable that straightens the front? Or are there lamellipodia/filopodia?

      We assume the reviewer is asking about the wavy edge outlined at 400 min after wounding (now Fig. 5A). As shown by Jacinto and colleagues in the first pupal wounding paper (JCB 2013), the actin cable forms quickly, within 15 minutes; much later actin protrusions extend from the leading edge to close the wound. This result is consistent with the wavy edge 400 min after wounding.

      248: The authors suggest an interesting hypothesis that mitochondria or ER could be pooled in fused cells. It would be nice to see some evidence: e.g. by labeling mitochondria and assessing where they are in syncytia versus mononucleated cells and whether they are concentrated around the wound edge.

      Although we don't think that exploring mitochondria or ER is central to this manuscript, we agree it would be an interesting question for the future.

      141-145 (Figure 4B and C) This example is not completely convincing. First, it is hard to see where the wound edge is. Second, it would be good to include an even later time point when the cell is clearly no longer at the wound edge.

      We have revised this figure, now Fig. 5B,C, to include a later image at 360 min after wounding healing, and this additional panel clarifies that the smaller cell leaves the wound edge. As noted in the text, the wound edge is indicated by the cell borders lacking p120ctn.

      Reviewer #3 (Public Review):

      Summary:

      White et al. described laser-induced wound healing of the Drosophila pupal notum. They found that the epithelial monolayer is dynamically induced to form syncytia by cell-cell fusion as an important part of repair. They reveal two processes: cell shrinking and border breakage that occur as part of syncytia formation. Expression of GFP in the cytoplasms of some epithelial cells reveals that cytoplasmic contents mix following injury and the GFP rapidly diffuses between cells. Using live imaging they observe that syncytia expand towards the wound, maintain their positions close to the leading edge, and apparently displace smaller cells. They propose that syncytia redistribute cellular components towards the wound facilitating repair and show that labelled actin becomes concentrated at the leading edge.

      Strengths:

      The manuscript is interesting and on an important and emerging topic of wound healing in a genetically tractable organism. The manuscript is very well written.

      Weaknesses:

      There are three major issues that the authors must address: 1. Is cell-cell fusion sufficient to enhance/facilitate wound healing? 2. Characterization of "border breakdown"; Is this phenomenon disassembly of apical junctions following membrane fusion? 3. Are cells really shrinking or is it only the apical domains that "shrink" as the cells join the syncytium.

      We thank the reviewer for recognizing the importance of this topic. Our responses to the specific weaknesses are below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Major Components:

      (1) For syncytia measurements the nuclei are labeled with histone-GFP which is expressed in all cell types. How do you know the nuclei within the cell junctions are epithelial and not another cell type, such as immune cells recruited to the injury site? It would be helpful to verify the number of nuclei per cell using an epithelial-specific nuclear marker as well. This could be via epithelial Gal4-specific expression of a UAS-nls-GFP.

      This is an interesting point. In response to the reviewer's question, we investigated by doing the converse experiment, labeling immune cells with hml-Gal4, UAS-GFP, and observing what they do after wounding (analyzing six wounded pupae). They do get recruited to the wound, but they remain either in the wound center or at the basal side of the leading edge. Because they are labeled with cytoplasmic GFP, we would be able to ascertain whether they fused with epithelial cells because they would share their GFP with epithelial cells in the epithelial plane, and they did not. Thus we are confident that the many syncytial nuclei are not derived from immune cells. Our live tracking throughout the manuscript, and specifically of GFP-labeled clones, also supports our interpretation that syncytial nuclei derive from epithelial cells.

      (2) The manuscript focuses on cell fusion, but other mechanisms of cell enlargement have been observed to occur during wound healing via endoreplication. To what extent do epithelial cells in pupae notum endocycle or endomitosis post injury? It is unclear if the increase in syncytia size during a 1-2hr period could also be due to endomitosis, which would also increase nuclear number.

      Since the first submission of this manuscript, we published our results demonstrating limited wound-induced endoreplication after this type of explosive laser injury to the pupal notum (White et al, 2024, PMID: 38495588). We chose to publish this work separately because we could not offer the same degree of depth for endoreplication as we could for fusion: our pupal notum injury model is extremely well-suited to analyzing cell fusion and wound closure by live imaging; however, it is not particularly well-suited for analyzing endoreplication in fixed tissue. With respect to reviewer's question about endomitosis -- i.e. nuclear divisions that are not accompanied by cell divisions -- even after many years we have not observed an endomitosis event, which would be visible by live imaging, whereas we frequently and easily observe mitosis of diploid cells.

      (3) One of the major conclusions of this study is that cell fusion is necessary to pool resources at the leading edge. Therefore it is critical that authors identify a mechanism to inhibit cell fusion to test this assumption.

      We now include new Fig. 4, an analysis of the role of Atg1 in promoting wound-induced fusion and wound closure. These results build on the finding of the Leptin lab (Kakanj et al, 2022) that autophagy genes are required for fusion. Our results are consistent with the model that syncytia speed wound closure.

      (4) There is evidence that myosin increases in endoreplicating cells during wound healing hence it is, maybe equally - if not more - probable that the increase in resources (here actin-GFP) at the leading edge is dependent on endoreplication instead of cell fusion.

      Some of the new data we provide for this manuscript is a correlation between cell size and distance traveled, showing that larger cells travel more within the wound (Fig. 4F,G). Endoreplication would certainly be expected to contribute to increasing cell size, and our published 2024 data indicates that there can be one extra S-phase induced by these types of wounds. Doubling the genome is not a significant contribution to cell size compared to the 10s of nuclei we observe in syncytia from fusion. Nevertheless, we do not claim that actin is the only important resource that can be pooled subcelluarly for the benefit of the cell; we use it only as a proof-of-principle. Finally, we discuss the work on myosin in wound-induced endoreplicating cells (Losick and Duhaime, 2021).

      Reviewer #3 (Recommendations For The Authors):

      Major comments

      (1) Can induction of epithelial fusion enhance wound healing?

      Different epithelial cell-cell fusion processes have been well-characterized: i) Trophoblast fusion in the placenta mediated by Syncytins. ii) Viral induced cell-cell fusion mediated by diverse viral glycoproteins (e.g. gp41 from HIV, Hemaglutinin from Influenza, GP from Ebola, and G glycoprotein from VSV). iii) Epidermal, myoepithelial, and other epithelial cell-cell fusion in C. elegans mediated by EFF-1 and AFF-1. iv) Cell-cell fusion in the eye lens (unknown fusogens). The authors may want to compare and discuss the temporal dynamics and intermediates observed in the diverse processes of epithelial cell-cell fusion with the characterization of syncytia formation during wound healing of the Drosophila pupal notum. Since some of these characterized cell-cell fusogens can fuse heterologous cells, including Drosophila S2 cells (Shilagardi et al., 2013; https://pubmed.ncbi.nlm.nih.gov/23470732/), the authors may consider expressing these fusogens in Drosophila pupal notum before, during and after injury. This could determine whether syncytia formation is sufficient to stimulate efficient wound healing.

      We thank the reviewer for the suggestion of comparing and discussing temporal dynamics and intermediates observed in the many types of epithelial fusion that are well understood. Regretfully, we do not think this article is the right venue for such a complex discussion, especially since we have little by way of comparison in our own wound-induced fusion data. As for overexpression of fusogens, it is an intriguing idea to force cell fusion with a heterologous fusogen such as EFF-1 and then investigate any resulting changes in wound healing. However, since half the cells within 70 µm of the wound already fuse even without a heterologous fusogen, it seems unlikely we could meaningfully increase the level of cell fusion unless we expressed the fusogen universally, forcing the fusion of nearly all the epithelial cells as well as other cells throughout the body that express pnr-Gal4. Because the overexpression of EFF-1 in C .elegans results in lethality (PMID: 26854231), a widespread induction of fusion would be expected to cause other types of physiological problems that would interfere with the interpretation of wound closure rates. Further, the conditional expression tools in Drosophila allow excellent spatial control, but temporal control is still somewhat low-resolution, so that we would have difficulty expressing EFF-1 before, during, and after wounding at times that would be relevant to understanding wound healing.

      (2) The phenomenon of "border breakdowns" described here is not clear. The authors are probably studying the disassembly of the apical junctions following the initiation of membrane fusion and pore expansion. This should be clarified by using membrane labels to directly observe membrane fusion. Researchers have used electron microscopy and membrane fluorescent probes to follow cell-cell fusion. For example, GPI-mCherry, FM4-64, lipid-modified-GFPs (e.g. PH-domain fluorescently labeled proteins) DiO, DiI, and many others. See for example: Markosyan et al., 2016; https://pubmed.ncbi.nlm.nih.gov/26730950/; Mohler et al., 1998; https://pubmed.ncbi.nlm.nih.gov/9768364/; Meng et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32668210/.

      We agree completely with the reviewer, that border breakdowns represent the disassembly of apical junctions following initiation of membrane fusion and pore expansion. Direct evidence for this order of events is found in the video stills of Figure 1 panel I and video S2, which show that cytoplasmic GFP is transferred to the fusion partner 14 minutes before there is a visible decrease in the apical adherens junction marker p120ctn. The reproducibility of this order of events is documented in Fig. 3: among 107 GFP-labeled cells, 30 of them first visibly shared GFP with a fusion partner, and then 11/30 displayed border breakdown, 16/30 displayed cell shrinking, and 3/30 did not fuse. This last category is consistent with a fusion pore that closed rather than expanded productively. Although we have obtained TEM images of wound-induced fusion pores, these are included in another manuscript currently in revision and so cannot be included here, and further these EM images do not shed light on border breakdown per se, as only live imaging can establish the relationship between border breakdown and pore formation (GFP-sharing).

      (3) The observation of cell shrinking may be misleading. The process the authors describe as "cell shrinking" may involve shrinking of the apical domain, maintaining the cell volume. To clarify this process, the authors may simultaneously label the apical and basolateral domains. It is possible that fusion pore formation occurs in the basolateral, apical, or both domains. The apical shrinking could reflect the migration of the apical junctions following fusion. A similar process has been described in epidermal and vulval cells of C. elegans and other nematodes (Mohler et al., 1998; https://pubmed.ncbi.nlm.nih.gov/9768364/; Sharma-Kishore et al., 1999; https://pubmed.ncbi.nlm.nih.gov/9895317/; Kolotuev and Podbilewicz 2008; https://pubmed.ncbi.nlm.nih.gov/18031720/).

      We thank the reviewer for pointing out these examples of cell fusion in nematodes, and we now compare our findings to Mohler et al, 1998. In Fig. 2D, we specifically investigated what happened to the cell volume of these shrinking cells, and we hope we have now clarified both the text and the annotations on the figure to make our findings more clear. In the X-Z plane, the entire cell volume of two shrinking cells is visible from cytoplasmic GFP labeling. For both cells, the cytoplasmic volume moves laterally into the neighboring syncytia, appearing to initiate the movement from the basal-most area of the cell so that 150 minutes after wounding, both cells have a reduced apical footprint and only a whisp of apically-oriented cytoplasm, with the remainder of the cytoplasm having moved into the syncytia. These images make it clear that fusion is occuring, and that when the apical area disappears the corresponding cytoplasm has also moved into the territory of the neighboring syncytium. In response to the reviewer's suggestion, we did try labeling basolateral domains, but the fluorescent proteins we examined are not restricted to the basolateral domain and are difficult to interpret.

      Minor comments

      (1) Lines 40-43. Repair of injuries has also been observed in non-proliferative syncytial epidermal cells and involves cell-cell fusogens. The authors may want to include this reference: Meng et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32668210/.

      We thank the reviewer for the suggestion, and we have included this reference in the Discussion paragraph about fusogens.

      (2) Lines 128-130. Is "Shrinking fusion" an "artefact"?

      The apical junction shrinks not the cell. I suggest following basolateral membranes to see whether the cell is indeed shrinking as it fuses. The authors may want to share whether the cell volume is maintained but spills into an existing syncytium; the apical junction shrinks because it disappears/disassembles (see also Major comment 3).

      As discussed in Major comment 3, we do provide evidence that the cell cytoplasm spills into an existing syncytium. Perhaps the reviewer finds the term "shrinking cell" to be misleading, as we all agree that the cell contents do not disappear. We have updated the manuscript to use the term "apical shrinking" throughout.

      (3) Lines 157-159. Are these small cells or instead they are small apical junctions? The interpretation should include basolateral domains of the small cells to determine their size! It is also possible that some small cells have fused with the syncytia but on the basolateral domain without apical junction disassembly.

      We appreciate the reviewer's rigor. As noted above, we were not able to analyze the basolateral domains of these cells. Because our all analyses are live-imaging videos, we are able to identify the cells are undergoing apical shrinking and clearly delineate those from stable diploid cells. We now realize that the term "small cells" is confusing and can be mixed up with apical shrinking. These cells are not "small" but normal sized, small only in comparison with the gigantic syncytia around them. We have removed the term "small" from this description.

      (4) Lines 204-206. Many genes required for myoblast fusion in Drosophila have been shown to play a role in different stages of cell-cell fusion. Do they play roles in epithelia fusion during wound closure in the pupal notum?. For example, actin polymerization? Dynamin? Ig-domain and integrin cell adhesion machineries?

      We now provide a new Fig. 4 that shows that the autophagy gene Atg1 reduces wound-induced cell fusion, as it does in larvae (Kakanj et al, 2022), and importantly these wounds close more slowly. We have not analyzed mutants in actin polymerization because we are confident they would interrupt many aspects of wound healing. The Galko lab has identified that integrins suppress wound-induced cell fusion in larval epidermis, but we have not tested these. We have a manuscript in revision demonstrating a requirement for Dynamin and other endocytosis genes in wound-induced fusion, and without dynamin-mediated fusion, these wounds close more slowly.

    1. eLife Assessment

      This study provides fundamental insights into the mechanisms of visual object categorization in primates through a scalable behavioral framework for assessing category learning and generalization in macaque monkeys. The evidence is compelling, based on extensive behavioral characterization, rigorous control experiments, and comprehensive comparisons with humans and computational models, although extending the model analyses to the secondary monkey experiments would further strengthen the conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      This study presents a systematic behavioral characterization of object classification abilities in macaque monkeys using a high-throughput touchscreen-based paradigm. The work shows that monkeys can learn and generalize many binary object classification rules, and compares their behavior with humans and computational models. A key finding is that monkey behavior is more closely aligned with visual deep neural networks, whereas human behavior is better captured by language-informed models. The study provides a useful benchmark for understanding visually grounded object categorization in nonhuman primates.

      Strengths:

      The study introduces a scalable and well-controlled behavioral paradigm for testing many object classification rules in macaques. The comparison across monkeys, humans, and computational models is a major strength and makes the work broadly relevant to visual neuroscience, comparative cognition, and computational modeling. The results provide an informative framework for distinguishing categorization based primarily on visual representations from categorization supported by semantic or language-based knowledge.

      Weaknesses:

      Some aspects of the interpretation would benefit from clarification. In particular, it remains somewhat unclear what stimulus-level factors drive image difficulty, how much training performance reflects general rule learning versus repeated reinforcement of specific images, and whether monkeys and humans apply the same category rules. The link between macaque IT representations and monkey behavior is also suggestive but not yet fully resolved, given the limited and separate neural dataset.

    3. Reviewer #2 (Public review):

      Summary:

      The paper tackles a very interesting question and provides a solid and systematic piece of data that may be useful for numerous NeuroAI works in the future. The question is how well can macaque monkeys with a "pretrained" visual system without human knowledge learn to categorize images based on different kinds of (sometimes arbitrary) category definitions. In general, I love the paper, and I think both the data and presentation of it are beautiful.

      Strengths:

      (1) The authors developed a scalable method for training and studying this behavior, and did an exhaustive evaluation of monkeys' behavior and learning process.

      (2) Beyond the behavior result, they performed extensive analysis and control experiments to isolate the cue monkeys are using to perform the categorization.

      (3) The extensive comparison of behavior with deep neural networks is also super interesting.

      (4) The authors performed a very careful examination of generalization behavior in monkeys, similar to standard practise in machine learning.

      (5) The presentation of the data is very beautiful and deliberately designed, kudos to the authors for their efforts!

      (6) I really enjoyed the further categorization task based on human knowledge, and the arbitrary rule task; this really pushes our understanding of the visual categorization and learning capability of monkeys.

      (7) The examination of *learning dynamics* in human vs monkey is also quite interesting, i.e., humans can "understand the rule" and learn much faster versus monkeys learning across a few days.

      Weaknesses:

      (1) Though all results are pretty cool, the organization of results, figures, and sections can be modified to flow even better.

      (2) Maybe provide DNN categorization and generalization results for the non-main monkey experiments (Figures 2,3), those comparisons can be really interesting too!

    4. Author response:

      We sincerely thank the editors and reviewers for their time and thoughtful feedback on our manuscript. The reviewers' constructive comments have been very helpful in guiding our revision plan. Below, we outline our plan.

      In response to Reviewer #1's comments on clarifying the factors that affect image difficulty and categorization rules, we will implement several revisions. First, to clarify what drives image difficulty, we will test whether image typicality within categories, quantified using methods such as Kramer et al. (2023; Sci Adv 9.17: eadd2981), can explain monkey categorization performance. Second, we will also examine whether performance on generalization images depended on their similarity to specific repeated images and on their category typicality. Third, to address whether monkeys and humans apply similar category rules, we will focus on images for which monkeys consistently made errors and examine whether these same images also yielded lower performance (i.e., longer reaction times) in humans.

      Reviewer #1 also raised an important question about how well macaque IT representations and behavior align. The IT categorization performance estimated in our manuscript is currently lower than monkey behavior, but this may reflect the limited number of recorded neurons. We will estimate ceiling IT performance as a function of neuron count and compare it with monkey and human behavior.

      In response to Reviewer #2's suggestion to enhance narrative flow, we will reorganize the text and adjust the ordering of certain figures and sections to ensure smoother transitions between findings and analyses. Specifically, we will more clearly state which parts of the manuscript establish monkeys' categorization ability and which parts compare their behavior with models or humans before performing a triangular comparison across all three.

      Regarding Reviewer #2's suggestion to test DNN performance on control experiments (non-natural stimuli, arbitrary categorization), we agree this is an excellent addition. We will perform these analyses and plan to report the results in the revised manuscript.

      We believe these revisions will substantially strengthen the manuscript and fully address the reviewers' feedback.

    1. eLife Assessment

      This useful study presents the first application of engineered NK-92 cell-derived extracellular vesicles displaying CD19 scFv for the treatment of systemic lupus erythematosus (SLE). The concept of using targeted extracellular vesicles as a "cell-free" alternative to CAR-T/CAR-NK therapies is good. However, the current results are incomplete and do not provide strong support for the experimental hypothesis, particularly with respect to EV purification, characterization, mechanistic validation, and adherence to current EV field standards. Several major concerns should be addressed to strengthen the translational relevance, reproducibility, and biological interpretation of the study.

    2. Reviewer #1 (Public review):

      Summary:

      This study constructed engineered NK-92 cell extracellular vesicles displaying CD19 single-chain variable fragment and evaluated their therapeutic efficacy in MRL/lpr mouse models of systemic lupus erythematosus, demonstrating that these vesicles could deplete B cells, alleviate lupus nephritis, and improve mouse survival. However, this strategy lacks significant innovation compared to existing research. The current results are not sufficient to provide strong support for the experimental hypotheses.

      Weaknesses:

      (1) This study proposes using engineered EVs displaying CD19 scFv to target B cells for SLE treatment. However, similar core therapeutic strategies have been reported in previous studies. For instance, recently, studies have reported engineered EVs for SLE therapy (J Control Release. 2025, 384:113886; Ann Rheum Dis. 2025, 84(11):1811-1821; J Nanobiotechnology. 2026, 24(1):203). Another research team from China also constructed engineered EVs displaying anti-CD19 scFv for SLE treatment, which is highly consistent with the present work in targeting strategy, delivery vehicle, and disease model (Mol Ther. 2026:S1525-0016(26)00080-8). Moreover, the human trial of allogeneic CD19-targeted CAR-NK therapy for SLE has been published (Lancet. 2026, 406(10522):2968-2979). This study has not made original improvements in therapeutic vectors, targeting modules, therapeutic mechanisms, and indications, and thus finds it difficult to meet the requirements of high-level journals for originality and novelty.

      (2) Numerous core experiments are missing, including the validation of CD19 scFv fusion protein expression on EVs, systematic characterization of engineered EVs, verification of EVs functions and therapeutic mechanisms, and in vitro and in vivo safety assessments. The available data are insufficient to support complete conclusions.

      (3) The stable expression of CD19 scFv on EVs should be further verified by Western blot or flow cytometry. The anchoring of CD19 scFv on the outer membrane surface of EVs must be confirmed. In addition, the loading capacity of CD19 scFv on exosomes should be quantified for the dosage selection in SLE treatment.

      (4) In vitro experiments are required to confirm the specific targeting ability of CD19 scFv-EVs to B cells and clarify the precise mechanism of B cell depletion, particularly whether it is mediated by effector molecules carried by exosomes such as perforin and granzyme B.

      (5) The key quality control parameters, such as the stability, purity, buoyant density, and particle/protein ratio of engineered exosomes, should be characterized and identified.

      (6) For the in vivo treatment experiments, the author needs to explain how the treatment dose of CD19scFv-EVs was determined in order to clarify the dose-effect relationship.

      (7) It is necessary to supplement with in vivo imaging and tissue distribution data to prove that the CD19 scFv-EVs can specifically accumulate in B-cell organs such as the spleen or lymph nodes.

      (8) The author needs to clarify the mechanism by which CD19 scFv-EVs reduce B cells in vivo and verify the caspase apoptosis pathway.

      (9) For the in vivo therapeutic experiments, the clinical first-line drugs and the free CD19scFv should be used to supplement the control group to highlight the advantages of the engineered EVs.

      (10) Safety assessment in this manuscript is completely absent. Routine toxicity examinations, including hepatic and renal function tests, routine blood tests, and histopathological analysis of major organs in mice, must be supplemented. In addition, the systemic inflammatory cytokine profile and anti-drug antibody levels should be determined to rule out critical safety risks such as cytokine release syndrome and immunogenicity. The authors only focused on alterations in B cells; the impacts of the treatment on T cell subsets, NK cells, and monocytes/macrophages should be further investigated.

    3. Reviewer #2 (Public review):

      Summary:

      Sun and colleagues report the development of an engineered extracellular vesicle platform derived from NK-92 cells that display an anti-CD19 single-chain variable fragment (scFv) on their surface via fusion with LAMP-2B (V-CD19-Exo). In an MRL/lpr mouse model of SLE, the authors demonstrate that intraperitoneal administration of V-CD19-Exo reduces splenic CD19+CD20+ B cells, attenuates proteinuria and lupus nephritis pathology, downregulates pro-inflammatory cytokines (IL-17A, IFN-γ) and autoantibodies (anti-dsDNA, ANA), and improves survival from approximately 25% to 80%. The authors propose that this "cell-free" targeted extracellular vesicle strategy offers advantages over conventional cell therapies, including lower immunogenicity, scalable production, and no requirement for lymphodepletion.

      The study addresses an important question in autoimmune disease therapeutics: how to achieve targeted B cell depletion while avoiding the complexities and safety risks associated with CAR-T/CAR-NK cell therapies. The concept is novel, and the initial in vivo efficacy data are encouraging. However, several significant limitations in experimental design, mechanistic depth, and evidence rigor temper the strength of the conclusions.

      Strengths:

      (1) Novel conceptual approach.

      The adaptation of CAR targeting principles to extracellular vesicles represents a creative and potentially impactful strategy. By displaying CD19 scFv on NK-92-derived vesicles, the authors successfully confer B cell-targeting capability while retaining the cytotoxic effector functions of the parental NK cells. This "cell-free" concept addresses genuine limitations of live cell therapies, including the need for lymphodepletion, risks of cytokine release syndrome, and manufacturing complexity.

      (2) Comprehensive in vivo efficacy readouts.

      The study evaluates therapeutic effects across multiple clinically relevant endpoints: B cell depletion (flow cytometry), renal function (proteinuria, UPCR), renal histopathology (HE staining with semi-quantitative scoring), systemic inflammation (IgE, IL-17A, IFN-γ), autoantibody production (anti-dsDNA, ANA), and survival. This multi-dimensional characterization strengthens the phenotypic evidence for efficacy.

      (3) Appropriate control groups.

      The inclusion of non-targeted NK92-Exo as a control allows attribution of the observed effects to CD19-mediated targeting rather than non-specific vesicle-associated activities.

      (4) Significant survival benefit.

      The improvement in survival from 25% to approximately 80% in V-CD19-Exo-treated mice is substantial and represents arguably the most compelling evidence for therapeutic potential in this model.

      Weaknesses:

      (1) Mechanism of B-cell reduction remains unclear.

      The manuscript reports a dramatic reduction in splenic CD19+CD20+ B cells (from 10.53% to 1.51%) following V-CD19-Exo treatment. However, the authors do not establish whether this results from direct cytotoxicity (e.g., perforin/granzyme-mediated killing, apoptosis induction) or from functional suppression/downregulation of CD19 expression. The authors speculate that the effect is likely mediated by cytotoxic proteins carried by NK-92-derived vesicles, but no data are provided to support this mechanism. Essential experiments would include the detection of apoptosis markers (Annexin V, activated caspase-3/7) in B cells, assessment of perforin/granzyme B content within V-CD19-Exo, or in vitro co-culture assays demonstrating direct B cell killing.

      (2) Small sample sizes.

      Most experimental endpoints were assessed with n=5 per group, which is marginal for detecting modest effect sizes and may amplify the influence of individual biological variation. While the survival study had n=10 per group, the main mechanistic and endpoint analyses would benefit from larger cohorts (n=8-10) to increase statistical power and robustness.

      (3) No dose-response or dosing optimization studies.

      All experiments used a single dose (10⁹ particles per injection) and a fixed schedule (twice weekly for three weeks). The absence of dose-response data leaves unclear whether the observed effects represent maximal efficacy or could be achieved with lower doses, and whether alternative dosing regimens could improve outcomes or reduce potential off-target effects.

      (4) Lack of safety assessment.

      The authors emphasize the theoretical safety advantages of extracellular vesicles over cell therapies, but no systematic safety evaluation is presented. Key missing data include: histopathological examination of non-target organs (liver, lung, heart, gastrointestinal tract), assessment of off-target immune activation (T cell responses, cytokine profiles beyond those measured), and evaluation of potential accumulation or toxicity with repeated dosing.

      (5) Incomplete characterization of the engineered vesicles beyond targeting.

      While the manuscript successfully demonstrates CD19scFv display and vesicle enrichment of exosomal markers, it does not characterize whether V-CD19-Exo retains the full spectrum of NK-92 effector molecules (perforin, granzymes, FasL, TRAIL, cytokines such as IFN-γ) at functional levels. Quantitative or semi-quantitative comparison of cargo between V-CD19-Exo and parental NK-92 cells or non-engineered NK92-Exo would help contextualize the observed in vivo effects.

      (6) Sex as a biological variable is not systematically addressed.

      The authors note in the Discussion that the same treatment showed more significant efficacy in male mice compared to females (data not shown), yet all main experiments were conducted exclusively in female mice. Given the strong sex bias in SLE epidemiology (approximately 9:1 female-to-male ratio) and potential differences in immune responses between sexes, this observation warrants systematic investigation rather than a footnote. Presenting the sex-differential data or alternatively, conducting adequately powered sex-stratified analyses would substantially strengthen the manuscript.

      (7) Translational claims are premature.

      The manuscript repeatedly emphasizes advantages over cell therapy (low immunogenicity, scalable production, no requirement for lymphodepletion) as if these are established properties of V-CD19-Exo. However, no experiments directly compare V-CD19-Exo to CAR-NK or CAR-T cells in terms of efficacy, immunogenicity, or safety. Similarly, claims of "scalable production" and "high batch-to-batch consistency" are not supported by any manufacturing or quality control data. These statements should be toned down or supported with empirical evidence.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript describes the development of engineered NK-92-derived extracellular vesicles (EVs) displaying CD19scFv for targeted treatment of systemic lupus erythematosus (SLE). Using a CD19scFv-LAMP2B fusion strategy, the authors generated EVs intended to selectively target pathogenic B cells in the MRL/lpr lupus mouse model. The study reports reductions in CD19⁺CD20⁺ B-cell populations, improvements in proteinuria and renal histopathology, decreased inflammatory cytokines and autoantibody levels, reduced splenomegaly, and improved survival outcomes following treatment. The work aims to position engineered EVs as a cell-free alternative to CAR-T/CAR-NK therapies for autoimmune disease treatment. While the concept is interesting and potentially translational, the study currently lacks sufficient methodological rigor, EV purification standards, mechanistic validation, and comprehensive characterization to fully support many of the claims presented.

      Strengths:

      (1) The study addresses an important unmet clinical need in systemic lupus erythematosus and explores an innovative cell-free therapeutic strategy.

      (2) The concept of combining CAR-like targeting approaches with engineered EVs is interesting and potentially translational.

      (3) The manuscript includes both in vitro and in vivo experiments, including functional renal assessments, immune profiling, histopathology, and survival studies.

      (4) The authors attempt to evaluate multiple disease-associated readouts, including proteinuria, cytokines, autoantibodies, splenomegaly, and survival outcomes, which strengthens the overall biological relevance of the work.

      (5) The use of engineered NK92-derived vesicles as a scalable alternative to CAR-NK therapy represents a potentially attractive therapeutic platform.

      (6) The in vivo therapeutic observations in the MRL/lpr lupus model are encouraging and warrant further mechanistic investigation.

      Weaknesses:

      (1) The EV isolation strategy is not sufficiently rigorous for defining the isolated particles as "exosomes" according to current International Society for Extracellular Vesicles/MISEV guidelines. The precipitation-based workflow without density gradient purification or SEC raises major concerns regarding EV purity and identity.

      (2) No direct validation was provided demonstrating successful surface localization or functional accessibility of CD19scFv on EV membranes.

      (3) The characterization of EVs is incomplete and insufficient. Additional positive/negative EV markers, purity metrics, and orthogonal characterization methods are required.

      (4) The absence of density gradient ultracentrifugation is particularly concerning, given the systemic injection of EV preparations into mice, as contaminating soluble factors and non-vesicular particles may contribute to the observed therapeutic effects.

      (5) The manuscript lacks adequate mechanistic studies explaining how engineered EVs mediate B-cell depletion or immune modulation.

      (6) The in vitro functional assays are weakly designed, particularly the use of A549 cells for evaluating CD19-targeted vesicle function.

      (7) Important methodological details are missing, including EV normalization strategies, flow cytometry gating controls, blinding procedures, and randomization approaches.

      (8) Several figures, particularly TEM and western blot images, are of low quality and difficult to interpret.

      (9) The study does not sufficiently exclude the possibility that observed therapeutic effects result from contaminating soluble immune mediators rather than EV-specific activity.

      (10) Broader immune profiling is lacking despite the systemic immune complexity of SLE.

      (11) The statistical analysis section includes tests that are not reflected in the Results section, creating concerns regarding data presentation and consistency.

      (12) Overall, while the concept is interesting, the manuscript currently falls short of the experimental rigor expected for high-impact translational EV studies.