10,000 Matching Annotations
  1. Jul 2026
    1. eLife Assessment

      This important study combines experimental evidence with computational modeling to suggest that neurons in the frontal eye field form two overlapping topographical maps, one for visual inputs and the other for eye movement directions. The experiments leverage the smooth cortex of the marmoset to provide solid evidence that two maps exist at different scales. The modelling work provides a potential mechanistic insight into how these patterns may be used for flexible visuomotor transformations.

    2. Reviewer #1 (Public review):

      Summary:

      Flexible natural behavior requires flexible sensory-motor mapping. In the visual domain, a visual stimulus at one location can guide a saccade toward another. How the receptive field (RF) and motor field (MF) properties of oculomotor structures support this flexibility is not known. Dotson and Reynolds address this question in the marmoset, using oblique Neuropixels penetrations across horizontal segments of the frontal eye field+, supplemented by electrical microstimulation. They report that visual RF and saccade MF vector angles each change smoothly with occasional abrupt jumps, that the two maps are organized as mosaics at distinct preferred spatial scales, and that a moiré interference pattern arising from a constrained spatial-scale mismatch between partially correlated mosaics reproduces the empirical distribution of RF-MF angular differences. They conclude that visuomotor flexibility is embedded in the geometry of mismatch and matches between visual and motor maps.

      Strengths:

      (1) The question is well-motivated. Sensory-motor mapping is known to be flexible, and asking whether the topographic relationship between the two maps itself supports that dissociation is a fresh reframing of a long-standing problem in oculomotor control.

      (2) FEF+ lies on the smooth marmoset cortical surface, which permits high-density horizontal sampling that would be difficult in the macaque arcuate sulcus, and oblique penetrations are a sensible way to track tuning across the surface. The dataset is substantial by the standards of the field (39 sites of high-density recordings across two animals, several thousand isolated units).

      (3) The data are thorough, and the convergence of three independent lines of evidence is the strongest feature of the paper. Unit recordings, electrical microstimulation, and two architecturally distinct generative models point to the same organization.

      (4) The central idea is conceptually novel. The proposal that flexibility can reside in the geometry of the maps, rather than only in time-varying activity, is original, and it generates concrete, testable predictions for tasks that require flexible visuomotor routing.

      Weaknesses:

      Major concerns

      (1) The analysis collapses each oblique penetration onto a single horizontal axis and pools angles across all cortical layers, treating cortical distance as purely tangential. Because the trajectory is angled, horizontal distance and depth are confounded, so some of the apparent RF-MF drift along a penetration could reflect a laminar transition, in addition to tangential mosaic structure.

      (2) RFs and MFs are estimated from the same free-viewing sessions in temporally adjacent epochs, leaving each measurement open to contamination by the other. Activity near a saccade can reflect peri-saccadic remapping rather than the stable retinotopic RF, and saccade-aligned activity following a recent flash can carry a residual visual component, given the long-lasting visual responses in FEF+ (>500 ms). Residual cross-contamination of this kind would tend to make RF and MF angles look more similar than they are, inflating the apparent local coupling and biasing the RF-MF difference distribution that the moiré model is fit to.

      (3) The paper claims that visual and saccade mosaics occupy distinct spatial scales, but the two preferred spatial frequencies are close, and the separation is summarized by overlapping "failed-test" bands rather than by a statistical test or confidence interval on the preferred frequency itself. The reliability of this separation is not established.

      (4) It is not clear whether the moiré model is a better model than the non-mosaic alternative. The moiré models are shown to be consistent with the data through failure to reject a Kolmogorov-Smirnov null, which is a weak form of evidence, and they are not benchmarked against a non-mosaic alternative or null model. The AM/NM convergence demonstrates architecture independence, but not that a mosaic organization is required.

      Minor concerns

      (1) The link from topography to behavioral flexibility (such as anti-saccades and other context-dependent transformations) is presented as a prediction but is not tested with any task manipulation. The work establishes an organizational principle and a plausible generative mechanism; whether that organization is actually exploited during flexible behavior remains open, and the framing should make this clear so the functional claim is not over-read.

      (2) It is unclear how relative depth (depth 0) is defined and how layer boundaries were assigned. The Methods mention common-average re-referencing for CSD and local field analyses, but no CSD or power-depth profile is shown to anchor the layer IV / depth 0 reference across penetrations.

      (3) The Discussion is brief relative to the strength of the claims. It would be helpful to address the concerns and alternative explanations above, where these cannot be fully resolved by the data.

    3. Reviewer #2 (Public review):

      Summary:

      The authors asked how the visuomotor system can keep visual selection and saccade targeting related but not rigidly coupled-the flexibility required for tasks like anti-saccades. They recorded visual receptive fields (RFs) and saccade motor fields (MFs) from individual neurons in marmoset frontal eye fields and adjacent premotor eye fields (FEF+) using obliquely inserted Neuropixels probes that traverse horizontal segments of the smooth marmoset cortex. They reported that visual and motor vector angles each changed smoothly with occasional abrupt jumps (a mosaic, rather than retinotopic, organization), that the two maps drifted with respect to one another, and that they were best described by mosaic-map models tuned to different preferred spatial frequencies. They then proposed that the offset in spatial scale, combined with partial shared structure between the maps, produced a moiré interference pattern, in which the distribution of local visual-motor angle differences matched the data.

      Strengths:

      (1) Unlike in the macaque brain, where FEF is buried in the arcuate sulcus, the marmoset cerebral cortex is lissencephalic (smooth). The authors used this feature to their great advantage and sampled horizontal mesoscale structure with oblique penetrations of ultra-high-density Neuropixels probes.

      (2) The mosaic framing was grounded in previous studies on direction maps in ferret V1 and MT, and the rate-of-change analysis (Supplementary Figure 2) plausibly reproduced the fracture-line phenomenon of those maps.

      (3) The authors confirmed the robustness of the finding by reaching the same conclusions using two architecturally distinct generators: the Fourier-based annulus model and the Gaussian-noise model.

      (4) They additionally provided an independent confirmation of the mosaic saccade-vector organization using electrical microstimulation. They reconciled the lower microstimulation spatial frequency by matching the spatial-averaging footprint (Supplementary Figure 6).

      (5) The Noise Mosaic (NM) model was used thoughtfully to decouple two properties that the Annulus Mosaic (AM) model confounds - spatial-scale offset and inter-map correlation - and to show that both an intermediate correlation (ρ ≈ 0.6) and a scale offset are required. Conceptually, a structural (topographic) substrate for visuomotor flexibility is a fresh alternative to the standard account in which flexibility lives entirely in time-varying activity on a single map.

      Weaknesses:

      The two claims here are not of the same strength. The first claim that RF and MF angles are organized as mosaics at distinct spatial scales was well supported. The second claim that a moiré interference pattern is the substrate for visuomotor flexibility was an inference rather than a direct observation, and several features of the design contribute to how strongly it can be held.

      First, the recordings were one-dimensional. Each oblique penetration yielded a line through the cortex, so the two-dimensional moiré pattern (Figure 4A) existed only in simulations; it was not reconstructed from the data. What the data provided was the marginal distribution of local angular differences, and the model was accepted when its simulated distribution was statistically indistinguishable from the empirical one. Matching a low-dimensional summary statistic is necessary but not strongly sufficient - multiple underlying architectures could produce similar 1-D difference distributions - so the moiré interpretation is best read as a plausible and parsimonious interpretation of the data rather than a confirmed mechanism.

      Second, the inferential logic needs to be strengthened. The "preferred" spatial frequencies were those at which a two-sample Kolmogorov-Smirnov test fails to reject equality between model and data. Failure to reject is not confirmation, and the width of the accepted band depends on statistical power, which depends on sample size. The authors did show that most parameter combinations were rejected, so the test did discriminate. That said, a continuous goodness-of-fit landscape with confidence intervals on the preferred SF, and a direct test that the visual and saccade preferred SFs differ, would better support the "distinct spatial scales" claim than visual inspection of two overlapping troughs.

      Third, the dataset was from two male marmosets, with 18 of 39 sites contributing to the core analyses, and the angular-difference distributions were pooled across penetrations and animals. This is standard for primate electrophysiology, but it means the spatial statistics were assumed stationary across the region, and individual variability in map layout was averaged over. The oblique-penetration geometry also added some uncertainty: the spatial-frequency estimates (in cycles/mm) are only as accurate as the reconstructed penetration angles (18.2{degree sign} {plus minus} 8.3{degree sign}), and angle error would propagate directly into the inferred scales.

      Fourth, while the authors suggested the moiré interference pattern can serve flexible routing for behaviours such as anti-saccades, this was not directly tested. Instead, they used free viewing and natural saccades, so the paper demonstrated a candidate substrate without testing whether behaviour employs it. This does not undercut the main findings, but readers should treat the functional narrative in the introduction and discussion as a set of predictions rather than results.

    1. eLife Assessment

      In this valuable study, Bhojappa et al. investigate the function of four septin-associated kinases (Elm1p, Gin4p, Hsl1p and Kcc4p) in controlling septin ring stability, actomyosin ring organization and constriction, and cytokinesis in budding yeast. The evidence supporting the conclusions is convincing. The results are a meaningful addition to previously published work and provide a clearer picture of a complex mechanism involving proteins of partially redundant functions.

    2. Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weakness:

      Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

    3. Reviewer #2 (Public review):

      Summary and strengths:

      In this study, Bhojappa et al. investigate the roles of the septin-associated kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and cytokinesis in budding yeast. Through quantitative analyses of kinase localization dynamics, septin organization, actomyosin ring (AMR) constriction, and cell morphology, the authors demonstrate that Elm1 and Gin4 play particularly important roles in maintaining proper septin architecture and cytokinetic progression. The work further identifies an interaction between the Gin4 KA1 domain and the Hof1 F-BAR domain and provides evidence that several cytokinesis-related functions of Gin4 are independent of its kinase activity. Artificial tethering approaches further supports that spatial organization at the bud neck is critical for the execution of septin-dependent cytokinetic processes. The authors combine live-cell imaging, quantitative analyses, biochemical interaction assays, and genetic perturbations to build a comprehensive framework for understanding how these kinases contribute to cytokinesis.

      Comments on revised version:

      The revised manuscript has been substantially improved. The authors have carefully addressed the concerns raised during review by providing additional experiments, analyses, quantifications, clarifications, and improved presentation of the data. The conclusions are now well supported by the experimental evidence. The study advances our understanding of the mechanisms linking septin organization to cytokinesis and will be of interest to researchers studying septins, cell division, and cytoskeletal regulation.

      Overall Assessment:

      This work provides valuable mechanistic insight into the coordination of septin organization and cytokinesis by septin-associated kinases. The experiments are carefully executed, the analyses are thorough, and the conclusions are supported by the data presented. The manuscript represents a useful contribution to the fields of cytokinesis and septin biology.

    4. Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p and Kcc4p. Elm1 and Gin4 show the strong knock-out phenotypes, whereas Hsl1p and Kcc4p show the weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Comments on revised version:

      I thank the authors for their thorough and thoughtful response to my review. The revised manuscript clearly reflects their efforts to provide rigorous and high-quality science. Addressing the concerns raised in the review required significant effort, but the improvements in the manuscript make it clear that the work was well worth it.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weaknesses:

      A control for non-specific binding is missing from the in vitro binding assay. The figures suffer sometimes from the very small text in the labels, which obscures understanding. Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

      We thank the reviewer for highlighting the strengths of our imaging pipelines and protein-protein interaction data. We have now included appropriate controls for the in vitro binding assays, which demonstrate that the observed interactions are specific (Fig. 2H). We have also revised all figures to improve clarity, including increasing font sizes to enhance visibility across panels. We agree that mapping the specific phosphorylation sites regulated by these septin kinases would provide valuable mechanistic insights. However, only a few studies have addressed this direction so far (Mortensen et al., 2002; Asano et al., 2006 and Marquardt et al., 2024) [1-3]. The current study focuses on the interplay among septin-associated kinases and their role in regulation of the actomyosin machinery (AMR). In this context, we highlight several key findings:

      (i) a molecular link between the septin kinase network and AMR through physical interaction between the KA1 domain of Gin4 and F-BAR domain of Hof1 (Fig. 2F-2H and S3H);

      (ii) a kinase-independent role for Gin4 in coordinating septin organization and AMR dynamics (Fig. 3A-3E, S3F & S3G, S3I and 4F-4H);

      (iii) a novel role for Hsl1 in regulating septins and the AMR downstream of Gin4 and Elm1, potentially through plasma-membrane binding (Fig. 4I-4K, 5A-5F, 8A-8G and S9H-S9J); and

      (iv) crosstalk between Gin4 and Hsl1 that is independent of their role in the morphogenetic checkpoint (Fig. 6A-6D and S6A-S6C).

      We have clarified this scope in the Discussion section and explicitly stated that mapping these phosphorylation sites will be an important direction for future work (Lines 623-626).

      Reviewer #2 (Public review):

      Summary:

      In this paper, Bhojappa et al. provide insights into the function of septin-related kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and actomyosin ring (AMR) structure and constriction. Their findings are both corroborative of and complementary to previous related studies.

      First, the authors provide a comparative analysis of the dynamic localization of these kinases at the bud neck, as well as a comparative analysis of defects in septin localization, splitting dynamics, AMR constriction rates, and cell morphology in kinase deficient cells. They find that septin localization and splitting kinetics, as well as AMR constriction rates, are significantly perturbed in elm1∆ and gin4∆ mutants but remain largely unaffected in hsl1∆ and kcc4∆. A similar trend is observed in terms of cell morphology and viability.

      Next, the authors focus on elm1∆ and gin4∆ cells, demonstrating that the residence time of the F-BAR protein Hof1 is significantly increased and defective in these mutants. Using yeast two-hybrid (Y2H) and in vitro binding assays, they show that the KA1 domain of Gin4 interacts with the F-BAR domain of Hof1, which may explain the cytokinesis-related functions of Elm1 and Gin4. Supporting this, they find that Gin4's role in septin localization, AMR constriction kinetics, and Hof1 bud neck localization is kinase-independent.

      The authors then conduct a series of artificial tethering experiments given their bud neck localization is mostly interdependent. They first demonstrate that artificially tethering Gin4 to the bud neck rescues the morphology defects of elm1∆ cells, with the strongest rescue observed when Gin4 was forced to interact with Hsl1-an effect that was also kinase-independent. Additionally, artificial tethering of Hsl1 to the bud neck restores the morphology of elm1∆ cells in a KA1 domain-dependent manner, suggesting that Hsl1 functions downstream of Elm1 to maintain normal cell morphology. Consistently, artificial tethering of Elm1 to the bud neck in gin4∆ cells rescues morphology defects, as well as defects in Myo1 localization and AMR constriction, but only in the presence of full-length Hsl1. The rescue fails in the absence of Hsl1 or when using a version of Hsl1 lacking the KA1 domain, which supports the role of Hsl1 downstream to Elm1 in cytokinesis.

      Strengths:

      Altogether, this study offers valuable insights into the mode of cytokinesis regulation mediated by the septin-related kinases, mainly Elm1, Gin4, and Hsl1, and would be an important contribution to the field of septins and cytokinesis after addressing current weaknesses.

      We thank the reviewer for the detailed summary and for highlighting the novel findings of our study.

      Weaknesses:

      (1) When assessing rescue of the elm1∆ phenotype, it needs to become clearer whether only morphology or also cytokinesis and septin organization are rescued.

      To clarify the extent of rescue observed in elm1Δ cells, we extended our analysis beyond morphological parameters by quantifying septin organization and AMR constriction dynamics. These analyses now show that artificial tethering partially restores septin organization and AMR constriction kinetics in elm1Δ cells in addition to improving cell morphology. These results are now described in detail in Fig. 5 and lines 373-406.

      (2) The quantification of the microscopy data does not always match up with the example images, and it's not always clear how the authors quantitatively analyzed their data.

      We revised the manuscript to clearly outline the quantification methods used for microscopy data analysis and specified the statistical tests, number of cells analyzed, and number of experimental replicates in the figure legends. We also clarified the criteria used for phenotype scoring and quantification in the Materials and Methods section. In addition, We replaced representative images where necessary to accurately reflect the quantified data throughout the revised manuscript.

      (3) The forced tethering data are key to the paper, but the lack of a summarizing table makes it difficult to grasp the full picture.

      We agree with the reviewer and have now included a new summary table (Table 1) that compiles the results of all artificial tethering experiments presented in this study, including the percentage of rescue in cellular morphology observed upon forced tethering of these kinases to the bud neck, thereby providing a clearer overview of these experiments.

      (4) Novel results and those confirming earlier results could be better distinguished.

      We have improved the overall clarity of the manuscript to distinguish novel findings from the results that corroborate previous studies, and have cited the appropriate literature throughout the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p, and Kcc4p. Elm1 and Gin4 show strong knock-out phenotypes, whereas Hsl1p and Kcc4p show weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Strengths:

      The combination of genetics, cell imaging, and biochemical characterization of proteinprotein interactions is attractive.

      We thank the reviewer for recognizing the significance of our findings and for the helpful suggestions.

      Weaknesses:

      (1) Imaging and data analysis is the main weakness of this manuscript. The authors must avoid manual counting and selection when easy analysis software can be used to limit bias. Instead of presenting unclear statistics of "percentage phenotypes", they need to define clear metrics to offer meaningful phenotype analysis.

      We agree that improving the quantitative rigour of the image analysis is essential for this study. Accordingly, we implemented a semi-automated image analysis workflow in the revised manuscript that defines reproducible metrics, such as aspect ratio, and reduces reliance on subjective phenotypic scoring. The inclusion of these parametric measurements enables clearer and more objective comparison of the rescued phenotypes.

      (2) This manuscript examines a very complex mechanism with four kinases of overlapping function using new data and existing literature. A clearer picture/model at the end of the manuscript that synthesizes the current knowledge would be beneficial:

      We incorporated a new representative model (Fig. 9) that integrates current knowledge in the field with our findings. This model highlights crosstalk among Elm1, Gin4, and Hsl1 as a key mechanism coordinating septin architectural transitions with AMR constriction during cytokinesis and is discussed in lines 520-536 of the revised manuscript.

      We sincerely thank all the reviewers for their insightful comments. We incorporated new results in Fig. S3A, S3B, S4A-S4F, 5A-5F, 6A-6D, S6A-S6C, 8E, 8F, S10B and 9, along with Table 1 summarizing the artificial tethering experiments in the revised manuscript. We believe that these revisions have improved the rigor of our analyses and enhanced the overall clarity of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 70: " abnormal cytokinesis defects": this language is redundant, either "abnormal" or "defects" would suffice:

      We thank the reviewer for identifying the redundant wording. We have rephrased the final paragraph of the Introduction section.

      Please refer to line numbers 86-97.

      (2) Currently, the final paragraph of the Introduction is an extensive, detailed summary of the results. This is unnecessary, as the Abstract and the Results sections summarize the results. Better would be a short statement of the questions addressed in the manuscript:

      We thank the reviewer for this suggestion. We have revised the final paragraph of the Introduction to remove the detailed summary of results and instead outline the key questions addressed in this study. The revised paragraph now emphasizes the knowledge gap regarding how septin-associated kinases regulate septin organization and coordinate cytokinesis. Please refer to line numbers 86-97.

      (3) In the Results, the wording of the heading associated with the first section is confusing. "and their defects during the cell cycle": it is not expected that the kinases themselves will have defects; defects may be observed in cells upon mutation of the kinases, but that is not clear from this wording:

      We thank the reviewer for this comment. The heading has been revised from “and their defects during the cell cycle” to “and defects associated with their deletions” for improved clarity.

      Please refer to line numbers 99-100.

      (4) Lines 107-109 (" they may act as molecular signals for the transition and a trigger for crosstalk of cytokinesis") are redundant with earlier lines 104-105 (" suggesting them to be a possible trigger for septin remodelling"):

      We have rephrased the text to improve clarity and avoid redundancy.

      Please refer to line number 123-125.

      (5) Lines 121-122: "Septin-associated kinases are believed to play an essential regulatory role": is the role essential, or is it regulatory? Since these kinases are not individually essential for cytokinesis, "essential" doesn't seem appropriate here:

      This has been corrected in the revised manuscript.

      (6) Figure 1 panels C and F: the font size is exceedingly small for the labels and should be greatly increased. This is true for multiple panels in Figures 2-5 and the Supplemental Figures as well:

      We thank the reviewer for highlighting this issue. We have significantly increased the font sizes of the text and the x- and y-axis labels across all figures in the manuscript.

      (7) The timing of mitotic spindle breakdown was used as a timepoint for comparison but it is not made clear in the manuscript how this was determined. Presumably, the mRuby2-Tub1 marker was visualized and the timepoint when the mitotic spindle separated into two discrete entities was called the "breakpoint" timepoint, but it would be important to better describe (and ideally show an example) of how that timepoint was determined. The only images I find with labelled tubulin are Tub1-GFP and these do not show spindle breakdown:

      We thank the reviewer for raising this point. We used GFP-Tub1 (pAFS125-GFPTUB1) and mRuby2-Tub1 (pHIS3p:mRuby2-Tub1+3′UTR::URA3) plasmids to visualize spindle dynamics across the cell cycle, and in both cases defined the spindle breakpoint as time zero. As suggested, we have now included time-lapse images of Cdc3-mCherry and GFP-Tub1 in wild-type cells in Fig. S2D to illustrate the spindle breakpoint event used for temporal alignment. Corresponding changes have also been made in the Materials and Methods section to explicitly describe this analysis.

      Please refer to line numbers 760-761.

      (8) Line 165: "F-BAR protein Hof1, which senses and induces membrane curvature" and lines 170-172 "the F-BAR protein Hof1, which is known to be associated with septin hourglass and transit to AMR during split ring trigger". It is awkward to introduce the same protein twice, in two different ways, within a few lines of each:

      We thank the reviewer for noting the redundant description. We have rephrased the text to introduce the F-BAR protein Hof1 in a more concise and streamlined manner while retaining the relevant functional information.

      Please refer to line numbers 206-208.

      (9) Throughout, it would be helpful to introduce more line breaks and organize the Results into smaller paragraphs:

      We thank the reviewer for this suggestion. We have revised the layout of the Results section and introduced additional line breaks to improve readability.

      (10) This is somewhat of a personal preference, but in the interest of transparency (and with the understanding that the 0.05 value is entirely arbitrary), would authors be willing to show actual P values in the figure panels rather than "**" or "ns", for example? Some readers may wish to apply a different standard of significance than 0.05, and not showing the P values makes this impossible. Furthermore, some readers (like this one) may interpret differently a P value of 0.044 versus "*" or 0.051 vs "ns":

      We agree with the reviewer’s suggestion. To improve the transparency of the quantitative analyses, we have modified the graphs across all main and supplementary figures to display the “actual p-values” and specified the corresponding statistical tests in the figure legends, with significance indicated by asterisks. An example is provided in the attached image showing the residence time of Inn1-mNG in wild-type and gin4Δ cells complemented with kinase-active and kinase-dead Gin4 constructs (Fig. 3E in the revised manuscript). For this analysis, significance was assessed using the nonparametric Kruskal-Wallis statistical test.

      (11) It is written that the authors "performed a Yeast Two Hybrid screen to find novel interacting partners at the bud neck" but I do not find evidence anywhere of a "screen" being performed, i.e., an unbiased search of many proteins to find a few interactors. Instead, it appears that the authors performed a two-hybrid assay to visualize interactions between a small, specific set of proteins. "Screen" should be replaced with "assay", as is currently the case in the Methods section. It is probably also valuable to point out in the text that using a yeast two-hybrid assay to assess interactions between yeast proteins has the caveat that any interactions observed could be indirect, as they may be "bridged" by endogenous yeast proteins.

      We thank the reviewer for raising this point. Our Yeast Two-Hybrid experiments were performed using a small subset of septin-associated proteins rather than as an unbiased screen. Accordingly, we have replaced the term “screen” with “assay” throughout the revised manuscript and in the Materials and Methods section.

      Please refer to line 241.

      We also agree with the limitations inherent to this assay and have now explicitly stated it in the revised manuscript, please see lines 248-255.

      (12) Panel 2H: I do not understand what the middle (as opposed to the top and bottom) blot segment represents. It is labelled "anti-HIS", like the one above it, but it is not associated with any molecular weight/ladder marker and I do not know what other species in the binding reaction in that lane would be recognized by the anti-HIS antibody. Perhaps the top segment is an "input" sample, and below it (in the middle segment) is what was bound to the beads. The figure legend is uninformative in this regard. Also, panel I in this figure is unnecessary to show, assuming that the bands shown in H are what I think they are. The blot makes the point without the need for quantification.

      As suggested by the reviewer, we have added molecular weight markers for each blot panel showing the input and bead-bound fractions. The figure legend has also been updated accordingly to clearly describe the different blot segments and experimental conditions.

      Please refer to figure legend 2H.

      In addition, as suggested by the reviewer, we have removed the quantification graph corresponding to the in-vitro binding assay from the revised manuscript.

      (13) The results in Figure 2H demonstrate that the Gin4-KA1 fragment is not non-specifically "sticky", because it does not bind GST alone, but there is no demonstration that binding by the Hof1 fragment is specific because there is no equivalent negative control for binding:

      We thank the reviewer for this suggestion. We repeated the in-vitro binding assay using 6His-bdSUMO as a negative control alongside 6His-bdSUMO-Gin4<sup>KA1</sup> to demonstrate binding specificity of the Hof1 fragment. The Hof1 N-terminal F-BAR fragment did not pull down the control 6His-bdSUMO fragment but specifically pulled down 6His-bdSUMO-Gin4<sup>KA1</sup> under identical experimental conditions (Fig. 2H), confirming the specificity of the interaction between Hof1 F-BAR domain and Gin4KA1.

      The corresponding text and results have been updated in the revised manuscript.

      Please refer to Fig. 2H and lines 259-263.

      (14) Line 257-258: "Localisation via Hsl1 is necessary to rescue the morphological defects exhibited by Δelm1 cells partially": what does "partially" refer to here? To the rescue, or the defects?

      The term “partially” refers to the extent of rescue. Specifically, elongated cell morphology was rescued in 63.75% of the elm1Δ cell population, rather than in all cells, upon artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP.

      (15) Lines 292-293: "can restore the morphological defects": this wording is unclear. "Restore" means "return to a former condition", which in this case would be normal cellular morphology, not defective cellular morphology. Similarly, see line 304: "While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering": presumably the proper localization was restored, not the mislocalization:

      In Lines 292-293, by phrase “can restore the morphological defects” was intended to indicate rescue of the elongated/clumped morphology associated with gin4Δ cells upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP. We have now rephrased this sentence as: “can rescue the elongated/clumped phenotype exhibited by gin4Δ cells”.

      Please refer to line numbers 481-482.

      Similarly, in Line 304, the statement “While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering” referred to rescue of the Myo1 mislocalization phenotype observed in gin4Δ cells. We have rephrased this sentence as: “However, Myo1-3xmCherry localization was restored to the bud neck upon artificial tethering of Elm1-GFP in gin4Δ cells”.

      Please refer to lines 506-508 in the revised manuscript.

      (16) Lines 295-296: "We find that Elm1-GFP tethering via Shs1-GBP, Bud4-GBP, and Hsl1-GBP" this should be "or", not "and":

      Thank you for pointing this out. We have now corrected the text accordingly.

      Please refer to line number 485.

      (17) Lines 311 and 312 refer to "Inn1-3xmCherry lifetime" but previously Inn1 residence time was measured. Since fluorescence lifetime is a distinct kind of measurement/ assay, it seems important to clarify here what kind of experimental data are being referred to:

      The term “Inn1-3xmCherry lifetime” was intended to describe the residence time of Inn1 at the cell division site, defined as the interval between the initial appearance of the Inn1 fluorescence signal and its complete disappearance during cytokinesis. For clarity and consistency, we have replaced the term “lifetime” with “residence time” in the revised manuscript.

      Please refer to the line numbers 511 and 512.

      (18) Discussion: "Septins are considered as the fourth cytoskeletal elements due to their extensive structural and functional diversity." This sentence is confusing, as it seems to imply that what defines a protein as being "cytoskeletal" is structural and functional diversity rather than anything to do with forming filaments, etc. The rest of this first paragraph of the Discussion also sounds like a summary of the background, is quite redundant with the Introduction, and should be shortened:

      We thank the reviewer for this suggestion. We have rewritten the first paragraph of the Discussion to improve clarity, reduce redundancy with the Introduction, and better emphasize the main findings of the study.

      Please refer to the line numbers 538-552 in the Discussion section.

      (19) Lines 365-366: "Gin4 and Hof1 are synthetic lethal": this should be revised to "gin4∆ and hof1∆ are synthetic lethal". This is a good place to point out that standard yeast nomenclature inserts the ∆ symbol after the gene name, not before (as is done in E. coli genetics, for example):

      We thank the reviewer for this suggestion. We have replaced “Gin4 and Hof1 are synthetic lethal” with “gin4Δ and hof1Δ are synthetic lethal” in the revised manuscript.

      Please refer to line number 576.

      In addition, we have now consistently placed the Δ symbol after the gene throughout the manuscript in accordance with standard yeast nomenclature.

      (20) Line 399: "We also performed an extensive GFP-GBP screens": again, here "screen" implies that a large collection of genes/proteins were assayed, perhaps in an unbiased way, which does not accurately portray what was actually done, which was an extensive tethering study using GFP-GBP:

      We thank the reviewer for this suggestion. We have replaced the term “GFP-GBP tethering screen” with “GFP-GBP tethering assay” throughout the revised manuscript. In addition, we have included a brief description of the specificity and functionality of GBP nanobody and its application in the GFP-GBP tethering strategy, extensively used in this study.

      Please refer to line numbers 326-336.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Analysis of morphological defects of elm1∆ does not directly reflect the defects in cytokinesis and septin organization. For example, the deletion of SWE1 rescues the morphological defects of elm1∆ cells but not the cytokinesis or septin mislocalization (Bouquin et al., 2000). Considering this, the authors should address whether artificial tethering of Hsl1 to Gin4 in elm1∆ cells rescues the septin and cytokinesis defects or just the morphology. Is the role of Hsl1 in cytokinesis dependent on its role in the morphogenesis checkpoint? Can the authors comment on how much the defects observed by Hsl1 tethering to the bud neck may be a result of bypassing the morphogenesis checkpoint?:

      We thank the reviewer for this important point. We performed time-lapse imaging of Cdc3-mCherry in strains where Gin4-GFP partially rescued the elongated phenotype of elm1Δ cells (63.75%) when tethered to the bud neck via Hsl1-GBP. Under these conditions, 64.29% of cells showed rescue of Cdc3-mCherry mislocalization, and Gin4 localization itself was restored to the bud neck in 57.85% of cells. We also examined Myo1-ymScarletI dynamics, while 76.92% of untethered elm1Δ cells displayed Myo1 mislocalization to the bud cortex, this was reduced to 11.36% upon Gin4-GFP tethering via Hsl1-GBP. Together, these results indicate that the morphological rescue observed in elm1Δ cells is accompanied by restoration of normal septin organization and AMR dynamics.

      Previous work (Bouquin et al., 2000) [4] showed that Swe1 deletion rescues cell elongation in elm1Δ cells without restoring septin organization. Consistent with this, we found that 60.66% of elm1Δ swe1Δ cells exhibited a round morphology, but tethering of Gin4-GFP to the bud neck via Hsl1-GBP in elm1Δ swe1Δ background did not further enhance morphological rescue. These results suggest that the rescue of cellular morphology observed in our tethering experiments may, atleast in part, depend on Hsl1-mediated regulation of the morphogenesis checkpoint.

      Importantly, despite the lack of additional morphological rescue, a clear restoration of septin localization was observed when Gin4-GFP was tethered to the bud neck via Hsl1-GBP in elm1Δ swe1Δ cells. Overall, these results suggest that while Hsl1-dependent morphogenesis checkpoint regulation may contribute to cell shape rescue, the restoration of septin organization is independent of Hsl1’s function in morphogenesis checkpoint and instead reflects a direct requirement for Gin4 and Hsl1 at the bud neck.

      Please refer to Figures 5, 6, and S6 of the revised manuscript for these additional data.

      (2) As suggested by the authors, the interaction of the Gin4-KA1 domain with the FBAR domain of Hof1 may explain the cytokinesis-related functions of Gin4. As an orthogonal approach, how does KA1 domain deletion of Gin4 affect cytokinesis and Hof1 bud neck localization?

      We thank the reviewer for this suggestion. We first examined the bud neck localization of Gin4-ka1Δ-GFP in comparison with full-length Gin4-GFP. We observed that the localization kinetics of Gin4-ka1Δ-GFP were significantly altered relative to the full-length protein, with reduced recruitment and earlier removal from the bud neck. We also analyzed the localization kinetics of Hof1-mNG in both gin4-ka1Δ and gin4Δ cells. Our results show that Hof1-mNG displays increased residence time and altered accumulation kinetics during cytokinesis in both genetic backgrounds. Thus, loss of the KA1 domain phenocopies loss of the full-length Gin4 and is consistent with disruption of the physical interaction between Gin4 and Hof1.

      Please refer to Figure S4 for these results in the revised manuscript.

      (3) The authors state that "Elm1 and Kcc4 were present at lower abundance at the bud neck (Fig S1A-D) compared to the higher abundance of Gin4 and Hsl1, as observed in their fluorescence intensities (Figures S1B-C) ". This is not evident in the figures. The authors should show a quantification of how they judged abundance at the bud neck:

      We thank the reviewer for this question. Quantification of septin kinase fluorescence intensity at the bud neck was performed using the established protocol for measuring protein accumulation kinetics described by Okada et. al. 2020 [5]. Time-lapse imaging for kinetic analysis of septin-associated kinases during bud emergence shown in Fig. S1A-S1D was carried out using a point-scanning confocal microscope with a 100×oilimmersion objective. Different laser intensities were required because the fluorescence signals of Kcc4 and Elm1 were comparatively weak and not readily detectable above cellular background under the imaging conditions used for Gin4 and Hsl1. The images shown in Fig. S1A-S1D are therefore displayed using differential contrast settings to facilitate visualization. We have now explicitly clarified this in the figure legend.

      A more direct comparison of septin kinase abundance at the bud neck is now provided in Fig. S1E-F, where localization kinetics during the HDR transition/septin remodelling stage were captured using a laser-scanning spinning-disk microscope under similar imaging conditions. We have additionally included raw fluorescence intensity profiles during the HDR transition to better illustrate the relative abundance of these kinases at the bud neck during cytokinesis.

      Please refer to Figures S1E and S1F in the revised manuscript for the updated images and quantitative analyses.

      (4) How did the authors determine G1 and M-phase in the experiments shown in Figures S1A-D? Can the authors mark these phases on the timelapse images? Also, How do the authors explain the different behaviour of Cdc3 in graphs S1A-D among different strains?

      We thank the reviewer for this comment. Cell cycle stages were initially inferred based on bud size, where kinase accumulation at the bud neck correspond to bud emergence (small bud, G1), and kinase disappearance coincided with septin splitting (large bud, M phase). However, we agree that accurate assignment of cell cycle stages would require specific cell cycle markers. To avoid confusion, we have removed the cellcycle-specific stage assignments from the Results section and describe the kinetics relative to t=0 (bud emergence).

      The differential dynamics observed in the Cdc3-mCherry kinetic profiles likely reflect heterogeneity within the the cellular population. To address this, we combined the normalized fluorescence intensity profiles of Cdc3-mCherry from the strains expressing GFP-tagged septin kinases during bud emergence and have included this data as reference (Author response image 1).

      Author response image 1.

      Plot showing spatiotemporal kinetics of Cdc3-mCherry in strains expressing either Elm1-GFP, or Gin4-GFP, or Hsl1-GFP, or Kcc4-GFP.

      (5) Forced tethering based experiments are one of the key sets of experiments for this work, but it is difficult to have a comprehensive understanding of all the data considering how large the data set is and how dispersed it is in the supplemental and main figures (Figures 3-4-5 and Figures S4-S5-S6). It would be helpful to provide a table summarizing the tested forced tethering’s and the phenotypic outcome in the tested yeast strains (Wt/mutant):

      We thank the reviewer for recognising the extensive dataset generated from the artificial tethering experiments and for suggesting the inclusion of a summary table. We have now added Table 1, which summarizes the proteins used in the GFP-GBP artificial tethering experiments, their genetic backgrounds, the total number of cells quantified across three independent replicates, and the phenotypic outcomes associated with bud neck tethering under each condition.

      Please refer to Table 1 and lines 342, 361, 451, 453, 456 and 492 in the revised manuscript.

      (6) In the introduction section, it would help the reader to provide more information on already known molecular roles of septin-associated kinases in septin organization and AMR. Later in the results section (i.e. Figure S1, S2, and S6), the authors extensively explain and show data that independently corroborate some earlier findings, which makes it difficult for the reader to distinguish novel findings from the repeated findings. I suggest shortening the text for the corroborative results, which will help to put more emphasis on their novel findings:

      We thank the reviewer for this suggestion. We have now included the canonical roles of these four septin-associated kinases in the Introduction section of the revised manuscript.

      Please refer to line numbers 79-85.

      We have also revised sections describing corroborative findings and explicitly cited previous studies wherever relevant in the Results section to better distinguish previously established observations from the novel findings presented in this work.

      Other minor comments:

      (1) In Figure 1D-1F, also show the data for hls1∆ and kcc4∆ - which are shown in S3AC in the current version:

      We have now included the Inn1-mNG residence time in hsl1Δ and kcc4Δ cells, alongside elm1Δ and gin4Δ in Fig. 1E.

      Please refer to Fig. 1E in the revised manuscript.

      (2) The authors should be more careful in interpreting their negative Y2H data in Figure 2F.

      We have now explicitly discussed the caveats and inherent limitations of the Yeast Two-Hybrid assay in the manuscript.

      Please refer to line numbers 248-255.

      (3) Please provide quantification for Figure 5A, Figure S6I:

      We have added quantitative analyses showing rescue of Cdc3-mCherry and Myo13xmCherry mislocalization in gin4Δ and gin4Δ hsl1Δ strains upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP.

      Please refer to Figures 8E, 8F, and S10B in the revised manuscript.

      (4) In Figure S6: label is missing "∆" in front of hsl1:

      We thank the reviewer for pointing out this error. We have corrected the labels accordingly.

      (5) As common consensus on yeast gene nomenclature, I suggest the use of "gene∆" instead of "∆gene":

      We thank the reviewer for this suggestion. We have now consistently placed the Δ symbol after deleted gene names throughout the manuscript in accordance with standard yeast nomenclature.

      (6) Lines (535-536): min(distribution) and max(distribution) in the formula is confusing. Clarify it or if possible use "minimum value", "maximum value" instead:

      We thank the reviewer for this suggestion. We have replaced the term “distribution” with “value” in the formula for protein accumulation kinetics analysis.

      Please refer to the updated formula in the Materials and Methods section (Lines 752753).

      (7) In line 86, "Dynamics of Septin-associated kinases and their defects during the cell cycle": Change the title as it is not clear what is meant by "their defects" given the discussed results under this title:

      We thank the reviewer for this suggestion. We have revised the section heading from “and their defects during the cell cycle” to “and defects associated with their deletions”.

      Please refer to line numbers 99-100.

      (8) On the Hof1-mNG image (Fig2C), show the line used for the line scan profile. Additionally, a similar line-scan profile could be useful in Figure S3G:

      As suggested by Reviewer 3, we removed the line-scan analysis from Figure 2 in the revised manuscript because Hof1 ring organization showed substantial heterogeneity across cells, making line-scan analysis difficult to interpret reliably.

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The % phenotype units are terrible. With no explanation, we do not really know whether they represent the percentage of cells that have a particular phenotype, or whether they correspond to a metric that measures some deviation between normal and extreme phenotypes. I would strongly recommend using precise quantitative metrics systematically (i.e. intensities, aspect ratios, division times, etc.) to properly quantify phenotypes:

      We thank the reviewer for suggesting the inclusion of precise quantitative metrics to assess phenotypic differences in the GFP-GBP tethering experiments. In the revised manuscript, we adopted quantification workflows that have been extensively validated and widely used in the literature, including those reported by Marquardt et al., 2024 (Fig. 7B and 7D) [2] from the Bi Lab. In response to the reviewer’s suggestion, we have now incorporated additional quantitative measurements, including cell area and aspect ratio (defined as the ratio of the cell’s major axis to the minor axis), for the experimental datasets presented in the manuscript.

      In the main figures, we now include aspect ratio quantification, while additional parameters are provided for the reviewer’s reference. We also quantified the fluorescence intensity of tethered proteins at the large bud neck and present these data together with the aspect ratio analysis in Figure 4 for elm1Δ cells in which Gin4GFP is artificially tethered to the bud neck via Hsl1-GBP. These quantitative analysis corroborates our qualitative observations and further strengthens our conclusions. Please refer to Figures 4D and 4E as representative examples.

      We have also changed the y-axis labels throughout the revised manuscript. For example, the y-axis in Fig. 7B is now labelled as “Cells exhibiting round morphology (%)”. Please refer to Fig. 4H, 4K, 7C, 8D, S5G, S6C, S7C, S9G and S9J for inclusion of aspect ratio quantification. Statistical analyses for the represented graphs were performed using Kruskal-Wallis nonparametric test, (N=3, n>150 cells/strain) (*: p<0.05, **: p<0.01, ****: p<0.0001, ns: p>0.05).

      Author response image 2.

      (2) Some quantitative analyses were performed manually where simple automated analysis should be performed to provide unbiased, accurate quantification:

      We fully agree with the reviewer that automated image analysis approaches, such as segmentation-based methods, are generally preferred for minimizing bias in morphological quantification. However, elm1Δ and gin4Δ cells exhibit severe phenotypes, including pronounced elongation and clumping, which makes reliable automated segmentation technically challenging for accurate quantification of parameters such as aspect ratio and cell area. For this reason, we used manual annotation for these analyses, as this approach enabled accurate delineation of individual cell and reliable measurements of morphological parameters such as cell area and size across the datasets despite being more time-consuming.

      Please refer lines 769-775 in the Materials and Methods section.

      (3) The "tethering" data also lack clear quantification. The authors should properly quantify the average intensity of Hsl1-GFP at the bud neck in each condition and correlate the results with cell aspect ratios or any other relevant parameters. For example, when comparing elm1null and elm1null Bud4-GBP with elm1null Kcc4-GBP, the visual impression is that as much Hsl1-GFP protein is recruited to the bud neck, whereas the phenotypes are dramatically different:

      We thank the reviewer for pointing this out. We have revised the image representation to facilitate clearer interpretation of the tethering experiments. In addition, we performed the key GFP-GBP tethering experiments using GBP-ymScarletI constructs, allowing direct visualisation of both the GFP-tagged protein and the GBP-tagged partner at the bud neck following tethering. We have included quantitative analyses of cellular morphology, raw fluorescence intensities of GFP-tagged proteins at the large bud neck, and the corresponding aspect ratio measurements for these updated datasets (Author response images 3, 4, 5). We have also included a summary table (Author response table 1) compiling these quantitative results for easier comparision.

      (4) I am very confused by Figure 3 which shows normal localization of Gin4-GFP in elm1null cells and seems to contradict other claims in the manuscript. This is very problematic for the interpretation of most of the "tethering" data:

      We understand the reviewer’s concern and have replaced the representative images of Gin4-GFP in elm1Δ cells in Figure 4B. Although Gin4-GFP is initially recruited to the presumptive bud neck during bud emergence in elm1Δ cells, it subsequently becomes mislocalized to the bud cortex during early cell cycle stages, resulting in reduced bud neck localization. As the cell cycle progresses, the Gin4-GFP signal at the bud neck decreases substantially in elm1Δ cells while remaining stable in wild-type cells until its departure prior to septin HDR remodelling (Fig. S5A-S5D).

      (5) Figure 6 is neither explained in the text nor in its legend. Could the authors explain the model and offer a comprehensive picture of the current knowledge?

      We have simplified the representative model to more clearly distinguish previously established knowledge from the findings presented in this study. Based on our results, We propose that Hsl1 functions both downstream of and in coordination with Elm1 and Gin4 to regulate septin stability and the timely execution of cytokinesis. Deletion of Elm1 disrupts the normal localization and crosstalk between Gin4 and Hsl1 at the bud neck, leading to septin mislocalization and misregulation of AMR dynamics, thereby revealing a previously uncharacterized role for Hsl1 in cytokinesis. The Results section has also been updated to reflect the revised model.

      Please refer to Figure 9 and lines 520-536.

      (6) Could the authors provide information about the double/triple mutant kinase phenotypes to clarify the overlap of functions among them?:

      Barral et al., 1999 [6] reported that individual deletions of Hsl1 and Gin4 result in mild cytokinetic defects, whereas deletion of Kcc4 does not produce any striking phenotype compared to wild-type cells. In contrast, the hsl1Δ gin4Δ kcc4Δ triple mutant remains viable but exhibits severe morphological abnormalities, including branched chains of elongated cells with defective cell separation. These mutants also display aberrant septin organization at the bud neck, characterized by irregular patch-like structures. Analysis of double mutants (hsl1Δ gin4Δ, gin4Δ kcc4Δ, and hsl1Δ kcc4Δ) revealed intermediate phenotypes between the corresponding single and triple mutants, with the hsl1Δ gin4Δ combination showing the strongest defects. Together, these findings suggest that the Nim1-related kinases function redundantly to regulate Swe1 activity and maintain septin architecture at the bud neck.

      Further supporting this model, Bouquin et al., 2000 [4] demonstrated that Elm1 operates independently of the Nim1-related kinases in controlling septin organization. The hsl1Δ gin4Δ kcc4Δ elm1Δ quadruple mutant exhibits severe septin localization defects and strong growth defects, in contrast to the elm1Δ single mutant, which primarily displays septin mislocalization from the bud neck to the bud cortex. These findings indicate that the combined activity of these kinases is essential for proper septin anchorage at the division plane and for assembly of the septin ring.

      Importantly, the progressively stronger phenotypes observed in double, triple and quadruple mutants also suggest that these kinases retain partially specialized functions at the bud neck. Consistent with this framework, our results support a model in which Nim1-related kinases function redundantly to regulate septin architecture and cytokinesis, likely through modulation of the AMR machinery. Because the localization of these kinases appears interdependent, as reported previously (Marquardt et al., 2020; Marquardt et al., 2024) [2,7] and corroborated by our findings, interpretation of mutant phenotypes remains complex and future studies will be required to delineate their individual contributions more precisely.

      Minor points:

      (1) Abbreviations are not defined in the manuscript.

      We have now expanded and defined all abbreviations throughout the manuscript.

      (2) Some of the writing in the figures is too small. Please make sure that a minimal size of letters/numbers is respected:

      We thank the reviewer for raising this issue. We have enlarged the figure labels and axis labels throughout the revised manuscript to improve readability.

      (3) Figure 2D. I am not sure that the line scans bring any useful information as the rings in mutant cells are quite heterogenous. This panel is also not cited in the text. Please make sure that every panel is cited at least once:

      We agree with the reviewer regarding the heterogeneity observed in the Hof1 ring organization in mutant cells and have therefore removed the line-scan analysis from the revised manuscript.

      (4) Knocked-out genes are written incorrectly. Please use the usual yeast nomenclature:

      We thank the reviewer for this suggestion. We have corrected the nomenclature for all deleted genes throughout the manuscript in accordance with standard yeast nomenclature.

      Author response image 3.

      Artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Gin4-GFP via Shs1-GBP-ymScarletI, Hsl1-GBP-ymScarletI, Bud4-GBPymScarletI and Bni5-GBP-ymScarletI in elm1Δ cells. DC*=Differential contrast. Scale bar5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple-comparison test (**: p<0.01, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=397, elm1Δ: n=408, elm1Δ-Shs1-GBPymScarletI: n=467, elm1Δ-Hsl1-GBP-ymScarletI: n=462, elm1Δ-Bud4-GBP-ymScarletI: n=401 and elm1Δ-Bni5-GBP-ymScarletI: n=317 cells). (C) Quantification of aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, n>170 cells/strain). (D) Graph depicting the raw fluorescence intensity of Gin4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=166, elm1Δ: n=177, elm1Δ-Shs1-GBP-YmScarletI: n=176, elm1Δ-Hsl1-GBPymScarletI: n=188, elm1Δ-Bud4-GBP-ymScarletI: n=185 and elm1Δ-Bni5-GBP-ymScarletI: n=151 cells).

      Author response image 4.

      Artificial tethering of Hsl1-GFP to the bud neck via septins or Nim1-related kinases rescues cellular morphology in elm1Δ cells. (A) Representative images showing the relocalization of Hsl1-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI and Gin4GBP-ymScarletI. Scale bar-5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=508, elm1Δ: n=418, elm1ΔShs1-GBP-ymScarletI: n=535 and elm1Δ-Gin4-GBP-ymScarletI: n=482 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (*: p<0.05, ****: p<0.0001), (N=3, n>165 cells/strain). (D) Quantification of raw fluorescence intensity of Hsl1-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=156, elm1Δ: n=166, elm1Δ-Shs1-GBP-ymScarletI: n=169 and elm1Δ-Gin4-GBP-ymScarletI: n=165 cells.

      Author response image 5.

      Targeted localization of Kcc4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Kcc4-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI, Hsl1-GBPymScarletI and Gin4-GBP-ymScarletI. Scale bar-5µm. (B) Quantitative analysis representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), oneway ANOVA Tukey’s multiple-comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=443, elm1Δ: n=489, elm1Δ-Shs1-GBP-ymScarletI: n=312, elm1Δ-Hsl1-GBP-ymScarletI: n=563 and elm1Δ-Gin4-GBP-ymScarletI: n=337 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (****: p<0.0001, ns: p>0.05), (N=3, n>165 cells/strain). (D) Quantification for the raw fluorescence intensity of Kcc4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (**: p<0.01, ***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=163, elm1Δ: n=168 elm1Δ-Shs1-GBP-ymScarletI: n=161, elm1Δ-Hsl1-GBP-ymScarletI: n=172 and elm1Δ-Gin4-GBP-ymScarletI: n=150 cells).

      Author response table 1.

      Summary table showing rescue of elongated morphology in elm1Δ cells upon forced recruitment of Nim1-related kinases via septins or its related kinases tagged with GBPymScarletI.

      Additional changes:

      The graph in Fig. S3D (revised preprint) has been updated to reflect a slight increase in the Chs2-mNG residence time in both the elm1Δ and gin4Δ strains, whereas our previous version indicated a delay only in the gin4Δ strain. Because the elm1Δ strain exhibited a more pronounced phenotype than the gin4Δ strain, we re-examined the analysis. The revised results show that the residence time of Chs2 during cytokinesis is modestly prolonged by approximately 2 minutes in both backgrounds. Accordingly, the graph and statistical analyses have been updated.

      References:

      (1) Mortensen, E.M., McDonald, H., Yates, J., and Kellogg, D.R. (2002). Cell Cycle-dependent Assembly of a Gin4-Septin Complex. Molecular Biology of the Cell 13, 2091-2105. 10.1091/mbc.01-10-0500.

      (2) Marquardt, J., Chen, X., and Bi, E. (2024). Reciprocal regulation by Elm1 and Gin4 controls septin hourglass assembly and remodeling. J Cell Biol 223. 10.1083/jcb.202308143.

      (3) Asano, S., Park, J.E., Yu, L.R., Zhou, M., Sakchaisri, K., Park, C.J., Kang, Y.H., Thorner, J., Veenstra, T.D., and Lee, K.S. (2006). Direct phosphorylation and activation of a Nim1-related kinase Gin4 by Elm1 in budding yeast. J Biol Chem 281, 2709027098. 10.1074/jbc.M601483200.

      (4) Bouquin, N., Barral, Y., Courbeyrette, R., Blondel, M., Snyder, M., and Mann, C. (2000). Regulation of cytokinesis by the Elm1 protein kinase in Saccharomyces cerevisiae. Journal of Cell Science 113, 1435-1445. 10.1242/jcs.113.8.1435.

      (5) Okada, H., MacTaggart, B., and Bi, E. (2021). Analysis of local protein accumulation kinetics by live-cell imaging in yeast systems. STAR Protoc 2, 100733. 10.1016/j.xpro.2021.100733.

      (6) Barral, Y., Parra, M., Bidlingmaier, S., and Snyder, M. (1999 Jan 15). Nim1-related kinases coordinate cell cycle progression with the organization of the peripheral cytoskeleton in yeast. Genes & Development 13. 10.1101/gad.13.2.176.

      (7) Marquardt, J., Yao, L.L., Okada, H., Svitkina, T., and Bi, E. (2020). The LKB1-like Kinase Elm1 Controls Septin Hourglass Assembly and Stability by Regulating Filament Pairing. Curr Biol 30, 2386-2394 e2384. 10.1016/j.cub.2020.04.035.

    1. eLife Assessment

      This useful study provides new insights into the liver-stage antigen LSA3, its export to erythrocytes, and its role in liver-stage development. While the functional importance of LSA3 is well demonstrated, the data underlying the conclusions regarding antibody specificity, liver-stage localization, and phenotype remain incomplete. A key strength of the study is the use of mosquito and humanized mouse models to access life-cycle stages that are rarely studied in most laboratories.

    2. Reviewer #1 (Public review):

      Summary:

      The extent P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood stage exported proteins tested in liver stages were not exported. An exception is LISP2 that is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages but in liver stages no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studied the localization of LSA3 in blood stages and used a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also showed that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 are thorough and mostly overcame adverse issues such as a cross-reactive antibody and negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      Weakness:

      The cross-reactivity of the antibody together with the co-infection strategy prevents reliable assessment of LSA3 localization in liver stages. Despite of this it seems LSA3 is not exported in liver stages and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. However, this was now done in a separate study focussing on blood stage parasites (PMID: 41135800).

      Due to the difficulty of working with humanised mice, it was not possible to refine some of the conclusions and questions remain:<br /> The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remain largely unknown. The co-infection used (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. It was also not possible to determine if the cross-reactive protein is expressed at that stage. While this might have been evident from the mixed WT/mutant infection (if all cells are positive for LSA3 there is cross-reaction; if about half of the cells are negative, there isn't) but assessing this failed.

      Significance:

      It is important information that LSA3 contributes to efficient liver stage development. However, neither LISP2 nor LSA3 seem to be exported in P. falciparum liver stages and can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

    3. Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood-stage LSA3 undergoes PEXEL processing and a portion of the protein is exported into the erythrocyte where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      (1) The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepesin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also reacts with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      (2) This study provides the first analysis of LSA3 exoerythrocytic function, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectible export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      Weaknesses:

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3-C reactivity at this stage could not be determined and the localization of LSA3 in the liver stage remains somewhat unclear.

      Comment on revisions:

      The authors thoughtfully addressed the reviewer comments.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      Comments on revised version.

      I appreciate the authors' efforts to revise the manuscript and to clarify several aspects of the study.

      However, I remain concerned that some conclusions extend beyond the data presented. In particular, the authors acknowledge in their rebuttal letter that antibody specificity in liver stages could not be validated and that cross-reactivity cannot be excluded. Consequently, the localization data shown in Figure 5 cannot currently be considered definitive evidence for liver stage localization of LSA3 itself.

      Similarly, the revised manuscript appropriately states that LSA3 was not detected beyond the PVM in liver stages and that export into the hepatocyte remains unresolved. Nevertheless, several statements continue to imply a role for liver stage protein export. At present, the possibility that a domain of LSA3 may face the host-cell side of the PVM remains speculative and is not supported by direct experimental evidence.

      The liver stage fitness phenotype is convincing and supports the conclusion that LSA3 contributes to normal liver stage development. However, the current data do not establish the developmental process affected nor connect the phenotype to export beyond the PVM.

      I therefore recommend that the manuscript consistently distinguish between (i) demonstrated export of LSA3 during blood stage infection and (ii) the unresolved localization and trafficking of LSA3 during liver stage infection. I would also encourage the authors to consider revising the title to better reflect the findings presented, as the current title may be interpreted as demonstrating liver stage export, which has not been shown.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The extent to which P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood-stage exported proteins tested in liver stages were not exported. An exception is LISP2, which is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, an incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum, it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages, but in liver stages, no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studies the localization of LSA3 in blood stages and uses a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also shows that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 is thorough and mostly overcame adverse issues such as a cross-reactive antibody and the negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable, as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      We thank the reviewer for their perspective on the strengths of the study.

      Weaknesses:

      The cross-reactivity of the antibody, together with the co-infection strategy, prevents reliable assessment of LSA3 localization in liver stages. Despite this, it seems LSA3 is not exported in liver stages, and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. Given that a previous publication found some inhibitory effect of LSA3 antibodies on blood stage growth, a comparison of the growth of the LSA3 disruption clones with the parent would have been very welcome and easy to do. At this point, LSA3 is one more of many proteins exported in blood stages for which the function remains unclear.

      It might be possible to refine some of the conclusions. The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remains largely unknown. The co-infection (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver, but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. The authors write that they do not know if the cross-reactive protein is expressed at that stage. But this should be immediately evident from the mixed WT/mutant infection. If all cells are positive for LSA3, there is a cross-reaction. If about half of the cells are negative, there isn't. In the latter case, the localization shown in the paper is indeed LSA3, and morphological differences between WT and LSA3 disruption could be assessed without additional experiments.

      We thank the reviewer for their comments. While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Significance:

      The conclusion from the paper that "our study presents just the second PEXEL protein so far identified as important for normal P. falciparum liver-stage development and confirms the hypothesized potential of exported proteins as malaria vaccine candidates" is partially misleading. Neither LISP2 nor LSA3 seems to be exported in P. falciparum liver stages, and we can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

      We thank the reviewer for the comment. We would like to emphasize the possibility that proteins localized at the PVM may be considered exported ‘if’ part or all of the protein (eg, a domain) faces the host cell lumen from the hepatocyte. We have not shown this to be the case for LSA3 or LISP2 but that possibility remains open. Nonetheless, LISP2 is exported (by P. berghei liver stages) and LSA3 is exported (by P. falciparum blood stages); both are exported proteins.

      Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif, but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al, PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood stage LSA3 undergoes PEXEL processing, and a portion of the protein is exported into the erythrocyte, where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepsin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also react with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage, where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      This study provides the first direct analysis of LSA3 function by reverse genetics, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage, but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectable export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      We thank the reviewer for their feedback regarding the strengths of the paper. 

      Weaknesses:

      A previous study reported that anti-LSA3 antibodies inhibit blood-stage growth, suggesting a role for LSA3 during erythrocyte infection. While the authors carefully evaluate the LSA3 mutant in mosquito and liver stages, the impact on blood stage fitness is not tested. While the knockout shows LSA3 is not essential in the blood stage, its importance during erythrocyte infection remains unclear.

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage, and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3C reactivity at this stage could not be determined, and the localization of LSA3 in the liver stage remains somewhat unclear.

      We thank the reviewer for their comments. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      We thank the reviewer for their comments.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      We understand the reviewer’s points. We agree that there is no evidence provided that LSA3 is targeted beyond the PVM; whether any of the protein faces the hepatocyte cytosol is unknown (and challenging to conduct) but this possibility remains plausible. LISP2- and LSA3deficient liver stages are less fit than parental controls and thus we stand by the conclusion that they are the two so far identified P. falciparum PEXEL proteins that are important for liver-stage development.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 163 says: "Altogether, this demonstrates that LSA3 is important but not critical for blood stage growth of P. falciparum": this is based on the cited Morita et al., 2017. However, previously LSA3 was considered dispensable based on a knock out in 3D7 (Maier et al., 2008; PMID: 18614010). Given that the authors generated a mutant for this work, it would be straightforward to test growth and clarify the importance of LSA3 in blood stages. If important, the analysis of the location and transport of LSA3 in blood stages would immediately become more relevant.  Maybe the data for this is already in the paper: the number of stage V gams was similar between mutant and control (Figure 4A). If this was calculated from the total number of asexual starting parasitemia, it includes blood stage growth, and it can be assumed that there is no growth defect in the mutant in the blood stages. If the number of stage 5 gams was calculated from the number of committed schizonts/rings, nothing can be said about blood stage growth, and asexual blood stage growth should be tested in specific experiments.

      We thank the reviewer for raising the function of LSA3 in blood stages and agree it was an obvious omission, though for good reason - a separate, collaborative study was underway. While this eLife preprint was in revision, our accompanying manuscript on the blood stage was published, showing the characterization of our NF54 DLSA3 mutant during blood-stage growth (PMID:41135800). The findings are now summarized and the citation included in the revised version of this preprint.

      Manuscript line 105: "although, notably, functional characterization of lsa3 deletion mutants has not yet been reported to confirm an important function": at least in blood stages, it was reported to be dispensable, see above. The corresponding study (Maier et al., 2008, PMID: 18614010) could be cited in that context. 

      The citation of PMID18614010 and 39913589 have now been added and we thank the reviewer.

      (2) Some questions central to the conclusions of this paper remain because it was unclear whether the serum did indeed detect LSA3 in the liver or not. It would be easy to check if all cells from the WT/Mutant mix experiment show LSA3 signal (this would mean it cross-reacts) or if only about half are positive (the mutants would be negative if there is no cross-reaction). This would be important to mention for Figure 5 because, at present, it is not known that what is labeled by the LSA3-C antibody in these images is (only) LSA3. 

      We thank the reviewer for this point and completely understand. We did check this via microscopy of liver sections co-infected with LSA3 mutant and control liver-stage parasites as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5-day post-infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (3) It is also unclear which parasites were imaged in Figure 5. The text of the results states that NF54 liver stages were used, but later: "As we employed a co-infection strategy to assess the essentiality of LSA3 versus NF54 in mice, we could not perform IFAs on individually infected mice in this study to validate the specificity of LSA3-C at the liver-stage". The legend says NF54 sporozoites on day 5 post-infection were used. I suspect it was a WT/mutant mix, in which case the above applies, and in the absence of cross-reactivity, half of the cells should be LSA3-C negative. If this is not the case, the localization in the liver becomes dubious.

      We apologize for the confusion and have corrected this. In Figure 5, we utilized liver sections from NF54-infected humanized mice that were stored at -80 C from a previously published study (McConville et al, PNAS 2024). Ideally, we would validate the specificity of LSA3 antibodies at the liver-stage using liver sections containing only DLSA3 parasites however the number of mice available was limited and the samples available to us also contained the Control line for qRT-PCR analyses (the co-infection strategy). As mentioned above, we couldn’t distinguish between these two strains by IFA at the time point analysed and this precluded us unequivocally validating the LSA3-C specificity in the liver-stage; however it cannot be excluded that the signal observed at the PVM is indeed LSA3. We are currently focusing research efforts on obtaining more humanised mice to answer this.

      Minor:

      (1) Introduction: Before the part on the PEXEL motifs, there are almost no references; please add references for all statements.

      We have added references.

      (2) Figure 1B is unclear regarding which part of the gene was deleted. The system used would permit a complete gene deletion, but the homology flanks seem to be within LSA3. If parts of the gene are left, the 75 kDa on the western blots might be a degradation product arising from both the truncated and the full-length protein. Please clarify in the sketch exactly where the homology flanks are, with respect to the start and stop of the gene. 

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (3) Line 161: Replace was with were.

      Corrected.

      (4) Figure 2, 224: Why do the authors think LSA3 must be in the luminal leaflet of the PVM as opposed to the outer leaflet of the plasma membrane?

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte.

      (5) Line 245: GFP core "derived from digestion of the reporter in the food vacuole, which confirmed it was secreted from the parasite". I wonder if the amount of GFP "core" really can be used as evidence for secretion, and its amount can be compared between experiments. Did the author quantify this for the full-length protein to get a proportion per sample?

      Use of GFP core to measure defects in P. falciparum GFP reporter secretion has been described previously (for example PMID:23387285 and 35906227). The comparison the reviewer asked for is an interesting and important question: however the control would be to compare the ratio of GFP core to uncleaved in the control lanes as well, which is not possible to do since the full-length protein is digested by plasmepsin V in the native PEXEL versions of the experiments (mLSA3-GFP in the first blot, Vehicle in the second blot) leaving no full-length protein to compare to. It stands to reason that inhibition of N-terminal processing results in less protein removal from the membrane (ER or COPII vesicle or PM) resulting in less secretion out of the parasite for retrograde transport to the food vacuole with cytostomal vacuoles (analogous to plasmepsin II). In the food vacuole, the chimeras are in normal cases digested by proteases back to the GFP core that is resistant to cleavage and evident as GFP core on the immunoblots (PMID:10775264 and 14709539 and 19055692 and 20130643). 

      (6) Figure 3 has the word plasmid in two lanes. In Figure 3E, amend the labelling of the blots.

      We apologize for the formatting error in converting the figures to PDF during the original submission and thank the reviewer for the suggestion. This has now been corrected.

      (7) Lines 266/271/284: "live IFAs", live immunofluorescence assay. Does this mean an antibody was given to living   parasites?

      The correct term is live microscopy and this has been corrected.

      (8) Does Figure 6A fit with the data in Figure 6B? It seems 6B has a milder phenotype than 6A.

      We thank the reviewer for the question. Yes the data directly correspond to each other and are represented in two ways: Panel A shows the qRT-PCR raw data for liver load of each parasite strain per humanized mouse using a scientific scale on the y-axis. Panel B shows that magnitude of the DLSA3 defect as a percentage of the total liver load per mouse:

      % total parasite liver load  = ( strain 1 or strain 2 liver load ) x100

      sum of strain 1 + strain 2 liver loads

      The intent of showing both data is to convey the correct magnitude of the difference in two ways to assist the reader in understanding the true defect, both are accurate and both are statistically significant. In revision we detected mislabelling of humanized mouse 2 and 3 in the original graphs that has now been corrected and we sincerely thank the reviewer for helping us identify this error.

      (9) Line 482: Please add references for this debate. 

      These have been added.

      Reviewer #2 (Recommendations for the authors):

      Major Comments: 

      (1) In general, the authors have taken care not to overstate conclusions from their study. Nonetheless, while not technically inaccurate, the title might misleadingly suggest LSA3 is exported in the liver stage (this was my initial impression on reading it until I looked at the data). I suggest the authors revise the title to avoid confusion by clarifying that export was only observed in the blood stage.

      We sincerely appreciate the reviewer’s point. As this article was posted as a preprint that has now been cited several times, we have carefully weighed the comment and in the end decided to retain the current title for the above reason.

      (2) While the ability to generate the ∆LSA3 parasites clearly shows that the protein is not essential in the blood stage, the impact on parasite fitness is never tested but simply assumed (for instance, in lines 163-164: "...this demonstrates that LSA3 is important...for blood-stage growth..."). Do the ∆LSA3 parasites have a fitness defect in the blood stage consistent with the previous GIA data that would support this claim? Since the rabbit anti-LSA3-C antibodies produced by Morita et al did not have GIA activity against the blood stage, it is possible that the GIA observed with the human and mouse antibodies might have been due to reactivity with a different protein. If ∆LSA3 does cause a fitness defect, it would be interesting to know if the endogenous GFP-tagged line, which alters protein trafficking/membrane association, also produces this effect.

      We agree with the reviewer and would like to clarify that this omission was not intended to create confusion but was by design, due to a separate collaborative study that was underway to address such questions. While this eLife preprint was in revision, our accompanying manuscript on characterising NF54 DLSA3 at the blood stage was published (PMID:41135800). The findings are now summarized and the citation included in the revised version of this eLife preprint. In sum, LSA3 is not critical for erythrocyte invasion but its deletion perturbs the rate and efficiency of merozoite invasion, at the step(s) of resealing of the PVM/host cell, resulting in aberrant accole forms that protrude from the infected erythrocyte.

      (2) Figure 1D: While the images are compelling and I don't doubt the claim that LSA3 is exported in the blood stage (also supported by the fractionation/Pk experiments), the authors should provide quantification of the difference in exported signal between the WT and ∆LSA3 parasites in these IFAs to rigorously support this conclusion. Also, please include details about how many independent experiments are represented by the microscopy data throughout the manuscript (Figures 1, 2, 3, and 5).

      We understand the reviewer’s request and wish to indicate that the export signal was absent in all cells infected with DLSA3 that was imaged. The microscopy performed was from n=2-3 experiments except for Figure 5 which was from n=1 humanized mouse per time point in which multiple EEFs from the liver were imaged. This has been indicated in the figure legends. 

      (3) Careful inspection of the z-series images in Figure 5A shows that most of the LSA3-C signal seen outside the PVM (beyond the boundary delineated by EXP1) is closely associated with DAPI puncta, suggesting these are merozoites. Together with the prominent gap in the EXP1 signal, this suggests the schizont has already ruptured. Thus, anti-LSA3-C signal beyond the PV seems best explained as coming from merozoites or other material released by PV rupture, not from export across the PVM, and this should be added to the text in place of comments about localization to PV extensions or potential export (lines 358-359, 422-423).

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have added the comment as requested.

      Minor Comments:

      (1) The authors may want to denote the disordered repeat region in the LSA3 schematic in Figure 1A that is mentioned in the text.

      We have added the residue boundaries of the predicted domain from AlphaFold into both the schematic and the text and included a link to the LSA3 pages in PlasmoDB and

      AlphaFold in the Methods section.

      (2) The authors use rabbit anti-LSA3-C antibodies previously generated by Morita et al. These polyclonal antibodies were raised against a recombinant C-terminal region of LSA3 (residues 750-1433), but the schematic in Figure 1A indicates the antibodies recognize a smaller region between residues 1154-1433. Please adjust the figure accordingly, or if this is not the same antiLSA3-C antibody reported by Morita, please provide details about its production.

      The figure is corrected.

      (3) The authors use Alphafold to identify a region of LSA3 with similarity to the substrate binding domain of DnaK, but the data is not shown. Please include the Alphafold prediction in supplementary figures and provide information about how the predicted structural homology was determined.

      We have added a link to the AlphaFold page for PF3D7_0220000 in the methods.

      (4) The schematic in Figure 1B indicates that the DHFR cassette was inserted at an internal site within the lsa3 gene. If this is the case, it seems possible that an N-terminal portion of the protein is still expressed, but I was unable to find details about the boundaries of the homology flanks to determine the precise insertion site. Please clarify the knockout strategy and indicate the specific insertion site.

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (5) Line 162: I think this should read "antibodies that react with LSA3 were...".

      Corrected.

      (6) Figure 1D: The merge with the transmitted light channel is missing for the third panel in the ∆LSA3 IFAs. Also, please define the scale bar length in the legend.

      Corrected.

      (7) Lines 744-746: The IFA fixation panel order description (top, bottom) in the Figure 2A legend is reversed from what is shown in the actual figure. Also, please define the scale bar length. 

      Corrected.

      (8) Lines 184-186: Since the fractionation/PK protection assays suggest most of LSA3 is in the PV, it would be interesting to know if the strong peripheral/PV signal observed in the PFA-fixed IFAs in Figure 2A is also present in the ∆LSA3 parasites, or is this non-specific? 

      Thank you for the suggestion. We agree this would be an interesting result to know but do not have the capacity at the present time.

      (9) Lines 219-225: It is unclear to me why these results are interpreted to suggest that the majority of LSA3 is peripherally associated with the luminal leaflet of the PVM. Wouldn't an integral membrane configuration in the PVM (with the C-terminus facing the host cytosol) or PPM (with the C-terminus facing the parasite cytosol) also account for the data? Adding a carbonate extraction would help clarify this point.

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM-associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte. If the question is whether LSA3 is an integral PVM protein, we agree that use of carbonate in the future would answer that question.

      (10) Figures 3D and E: There are some problems with some of the text wrapping in these panels.

      We apologise, this was a formatting issue as the manuscript was converted to PDF.

      We have corrected this error.

      (11) Line 422-423: In fact, the Z-sections shown in Figure 5 appear to indicate that the LSA3-C signal is predominantly located within the parasite, not at the PVM.

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have corrected the final conclusion to be more accommodating of this.

      (12) Lines 468-470: Since cross reactivity of anti-LSA3-C is substantial in the blood stage but was not defined in the liver stage by analysis of unmixed infections, how do the authors know that they were not observing ∆LSA3 parasites in their IFAs? I think what they mean here is that parasites lacking anti-LSA3-C reactivity were not observed, which is an important distinction.

      The reviewer is correct and this has been corrected.

      (13) Lines 478-479: The authors should also mention that the P. berghei PEXEL proteins evaluated in Fougere et al are exported in the blood stage, similar to LSA3. Moreover, other studies have shown something similar for additional endogenous PEXEL proteins or reporters in P. berghei (PMIDs 22329949, 26347246, 34956312).

      We have added the additional text regarding export into the infected erythrocyte and the reference to IBIS1.

      (14) Line 491: The data here don't support that LSA3 is "required" for liver stage development, only that it is important to it. Since the authors have not defined the cross-reactivity of anti-LSA3C in unmixed infections, it is not clear that ∆LSA3 parasites are arrested early in the liver stage, only that they show a reduced number of genome copies relative to the parental control. 

      We have amended the sentence to “required for normal liver stage development”.

      (15) Line 530: I think NGF54 should be NF54.

      Corrected.

      Reviewer #3 (Recommendations for the authors):

      (1) Antibody specificity in liver stage IFA experiments:

      The specificity of the anti-LSA3 antiserum (LSA3-C) used in liver stage IFA is not fully convincing. While the KO parasites were used effectively to validate specificity in blood stages, the same is not true for liver stages. 

      (a) It is essential to repeat IFA with ΔLSA3 parasites in liver stage infections to determine whether the observed PVM staining is truly specific.

      We appreciate the reviewer’s point, however at a cost of over $5000 per humanized mouse, we do not have the capacity to conduct this experiment at the present time. We highlight that, as the blood stage IFAs confirmed the specificity of LSA3-C for LSA3, the possibility remains open that LSA3 is specifically recognized at the PVM.

      (b) If the antibody is the same polyclonal serum used in Morita et al. (2017), why did the authors not employ a monoclonal antibody, which they presumably have access to and which would provide greater specificity? 

      We have included new data confirming that LSA3 is exported using LSA3-T, in addition to LSA3-C.

      (c) Given that rabbit antisera often show non-specific staining at the PVM in liver stage parasites, co-localization with PVM markers is not sufficient. Inclusion of the ΔLSA3 parasites in liver stage IFA is critical. It will also show whether there is any cross-reaction of the antiserum in liver stage parasites, as seen by IFA for blood stage parasites. 

      We thank the reviewer for their feedback.

      (d) To validate the serum further, the authors should infect HC-04 cells in vitro with GFP-LSA3 parasites and stain with LSA3-C to confirm overlap between the tagged protein and the antibody signal.

      We thank the reviewer for their feedback.

      (e) For higher-resolution co-localization, expansion microscopy - now commonly used even in malaria research - would substantially improve the analysis. 

      We thank the reviewer for their feedback.

      (2) The localization of LSA3 in this study differs notably from Morita et al. 2017, who reported localization to dense granules in merozoites and staining in ring-stage parasites at the PVM. 

      (a) The authors confirm DG localization, but they do not examine ring-stage parasites. They should include the IFA of ring stages to clarify whether they can replicate the previous findings.

      We thank the reviewer for their feedback.

      (b) Additionally, the differences in Western blot banding patterns between the two studies should be addressed. Do the authors have an explanation for these discrepancies? 

      We thank the reviewer for their feedback.

      (3) The authors report a ~40% reduction in liver parasite load using qPCR, which is statistically significant. However, this phenotype is modest and should not be interpreted as showing that LSA3 is essential.

      (a) Please avoid terms like "required" or "essential" and instead describe the protein as "contributing to normal development" or "influencing fitness."

      We have used the term “required for normal liver stage development”.

      (b) Since the authors generated liver sections, they should take advantage of these to quantify the number and size of liver stage parasites, which would help determine whether the phenotype reflects fewer infected cells or reduced parasite growth.

      We did check this via microscopy of liver sections, but all mice were co-infected with LSA3 mutant and control liver-stage parasites, as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5 day post infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (c) It would also be valuable to include IFA from singly infected ΔLSA3 livers (rather than co-infected), and possibly at earlier timepoints, to identify the developmental window affected.

      We agree it would be valuable.

      (4) The manuscript suggests that LSA3 may be exported beyond the PVM into the hepatocyte, based on a small number of peripheral puncta.

      (a) This claim is not convincingly supported by the data. The punctate signals shown in Figure 5 are weak and may rather reflect PVM extensions or TVN. In fact, one punctum even overlaps with the DAPI signal (figure 5, middle panel), which raises further doubt about the localization.

      We appreciate the reviewer’s careful eye and caution and have added the comment regarding DAPI.

      (b) Given the lack of KO controls in these liver stage IFAs, the authors should not describe LSA3 as "exported beyond the PVM". The language should be revised to reflect that the protein localizes predominantly to the PVM, and any extra-PVM signal remains unconfirmed and could be non-specific. 

      (c) This is especially important given the well-known tendency of rabbit antisera to produce background PVM staining in liver stage parasites. 

      Corrected.

      (e) In an earlier report (McConville et al, 2024, PNAS), they clearly state that LSA3 is NOT exported beyond the PVM. Actually, the staining in the previous report looks quite different from the images provided for Figure 5. The authors might wish to comment on this. 

      We thank the reviewer for their feedback.

      Minor comments:

      In some sections, the manuscript uses "exported" to refer to trafficking to the PVM. This terminology should be used more carefully and consistently, since "export" often implies translocation into the host cytosol

      We understand that export involves a protein localizing within the host cell and so protrusion through the PVM may also be considered exported, however, we have not confirmed this for LSA3 in liver stages.

    1. eLife Assessment

      This study offers an important set of literature-based parameter ranges for calibrating cardiac electromechanical models. The key scientific claims are largely supported by convincing evidence and carefully-documented methodologies, but a few specific issues remain only partially supported, in particular relating to the locking-free formulation, reporting of adequate details regarding the electrophysiological model (Eikonal approach), and the fact that physiological accuracy is predominantly achieved in a single chamber (the LV).

    2. Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains and other biomarkers. The overarching aim of the study is to give credibility to the underlying computational electromechanics framework as a step towards the cardiac Digital Twin vision.

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study relies on several simplified modeling assumptions that do not reflect the current state-of-the-art:

      a) Isotropic scaling of the ventricular mesh to generate an unloaded reference geometry.

      b) Simplified afterload and preload models that do not consistently capture the full range of physiological responses.

      c) Simplified epicardial boundary conditions.

      These limitations are appropriately acknowledged and discussed by the authors in a dedicated Limitations section.

      (2) Numerical Framework:

      a) The numerical framework used for the mechanical part of the model may be susceptible to locking effects that could contribute to artificially stiff and less contractile behavior. This is indicated by a ten-fold scaling of the peak active contractile force parameter relative to literature values and notable sensitivity of the model to the tissue compressibility parameter. While - as acknowledged by the authors - part of this can be attributed to simplified modeling choices, other comparable studies have not reported similar issues.

      b) The human electrophysiology model is not described in enough detail. Currently, it is not mentioned in the manuscript that an Eikonal model was used to compute activation times on the endocardial surface to be robust against coarse mesh resolutions.

      (3) Geometrical model and digital twin: The model presented combines anatomical data, electrical measurements, and physiological reference values from different individuals or population averages, rather than being derived from a single patient. The authors have appropriately moderated their claims in the revision, framing the work as a step towards the cardiac Digital Twin vision rather than asserting that the model itself constitutes a digital twin.

      (4) Calibration procedure: The description of the calibration procedure has been substantially improved in the revision. The authors now provide explicit rationale for each calibration step and clarify that the procedure targets multiple physiological biomarkers in sequence. Verification that the calibrated model produces physiological cellular dynamics, including intracellular calcium transients, is now provided. The revised manuscript also shows the simulated electrocardiogram alongside population reference ranges, which partially addresses the question of calibration quality. However, a direct comparison of the simulated electrocardiogram with the individual measured signal used for calibration is not provided, which would give a more stringent and direct assessment of how well that specific calibration target was achieved.

      Comments on revised version.

      The revision represents a genuine improvement. The calibration procedure is now more transparently described, physiological cellular dynamics are verified, the digital twin framing has been appropriately moderated to reflect a step towards that vision rather than a claim of having achieved it, and an expanded limitations section identifies where the framework falls short.

      Several of the concerns raised in the first round have been addressed, but some issues remain:<br /> The model still requires a ten-fold scaling of the peak active contractile force relative to literature values, and the imbalance between left and right ventricular output persists, along with non-physiological right ventricular pressures and ejection fraction. These are partly fundamental limitations of the current modelling approach that may not be fully resolvable within the scope of this paper, and the authors are to be credited for acknowledging them. However, they do constrain the conclusions that can be drawn about the credibility of the framework for reproducing healthy cardiac physiology.

      The population-averaged reference dataset and the calibration and validation framework remain contributions of value to the community, and the revised limitations section adds useful transparency about the current state of the art.

    3. Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Comments on revised version.

      The authors have addressed my prior comments

    4. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the reviewers for their efforts in reviewing our manuscript. We highlight that the key contributions of our paper are to provide a framework for calibration and validation of high-fidelity cardiac electromechanical models based on a diverse compilation of clinical datasets, and that we provide one example of such an evaluation of our own baseline electromechanical model. The comments raised by the reviewers were chiefly focused on the second goal, which is our specific model and the outcomes of the evaluation process, rather than on the evaluation framework itself. As such, we have made improvements to our model implementation and to provide additional confidence in our specific modelling framework through this review process. Specifically, we have strengthened the verification component of this evaluation, provided additional quantitative measures, and included a more in-depth discussion of the remaining limitations in our modelling framework. We hope that the updated version of the manuscript and our efforts to improve it are well-received by our reviewers and editors, as well as by the modelling and simulation community at large.

      eLife Assessment

      This is a potentially important study that explores the relevant range of parameter values for calibration and validation of cardiac electromechanics in ventricular models. Although much of the work presented is solid, the evidence provided to support the authors' key scientific claims is incomplete, especially as it relates to the emphasis on standardized validation and verification approaches. Notably, the level of model personalization presented in this work falls short of the threshold for what could reasonably be called a "digital twin", even by the relatively relaxed standards that have emerged in computational physiology and related fields in recent years.

      We appreciate the eLife assessment for identifying the potential importance of our study. Regarding the threshold for 'digital twin', we note that a cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning, which is a goal that to our knowledge no published electromechanical study has simultaneously fulfilled. It is for this reason that the community refers to the 'digital twin vision' rather than its realisation, and we adopt this framing consistently throughout the manuscript.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers within a single study. In our revision, we have clarified that the model evaluation presented here is an example application of that framework, which was designed not to certify a model as complete, but to provide a transparent audit of current capability that identifies where confidence is established and where further development is needed. We have updated the title and language throughout the manuscript to reflect this framing consistently.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction, and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains, and other biomarkers. The overarching aim of the study is to give "credibility to the underlying computational electromechanics framework" and to "pave the way towards credible cardiac electromechanical Digital Twins."

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study employs simplified modeling assumptions that are not state-of-the-art, e.g.,

      (a) Isotropic scaling of the mesh to generate an unloaded reference geometry.

      (b) Simple afterload and preload models that fail to produce physiological results.

      (c) Simplified epicardial boundary conditions.

      While our model was able to broadly achieve physiological behaviour based on the calibration and validation datasets, it also contains several simplifications that can be expanded with more sophisticated techniques to allow explorations in specific areas. We have added a dedicated Limitations subsection to the Discussion section of the manuscript to address these and to provide references to relevant studies.

      (2) Numerical Framework:

      (a) The mesh resolution and/or the numerical framework used for the mechanical part appears to suffer from known numerical artifacts (locking effects), leading to overly stiff or inaccurate behavior in finite element analysis. This results in an artificially stiff response to deformation, which is compensated by setting active contraction to ten times the value reported in the literature. The authors attribute this to limitations in using ex vivo tissue measurements to represent in vivo function, although similar issues were not observed in previous works.

      We thank the reviewer for raising this point and have investigated it carefully. We have added a verification section as well as discussions to the manuscript to more comprehensively discuss this point. In short, through various tests against benchmark (Land) and comparing stress-strain curves in cube simulations with the same mesh resolution, we could not identify evidence of volumetric locking effects. The elevation in contractile force was also necessary in a simplified ellipsoid version of the model in a previous publication [ref 7, Levrero-Florencio, et al. 2020]. We note that these benchmarks were performed in the incompressible transversely isotropic regime; whether analogous locking effects exist in the dynamic orthotropic active contraction framework used in the full biventricular simulations remains an open question, which we have identified as a priority for future benchmarking, for example against the Arostica et al. 2025 benchmark.

      We agree that the explanation of this as ex vivo vs in vivo difference in contractile force is too simple, and other contributing factors are better understood through comparison with similar studies in the field. Strocchi et al. (2023) used a four-chamber model with explicit atrial mechanics, and in her history matching varied Tref within +- 33-55% of a reference value of 120-150 kPa, targeting a peak active tension of 160 +- 15 kPa, which was a considerably more modest adjustment than applied here, likely reflecting differences in model geometry, pericardial constraint, and circulatory model between the two studies. Gerach et al. (2021, Mathematics) applied manual parameter adjustments informed by in vivo active tension measurements of 120 – 150 kPa and achieved ejection fractions of approximately 63%; however, they reported that systolic pressures in both ventricles were too high for a healthy heart, and similarly reported elevated peak ejection rates compared to MRI measurements, a difficulty we also encountered. Notably, Gerach et al., report that atrial contraction contributes approximately 11-13% of end-diastolic volume, which in a biventricular-only model would directly reduce the achievable LVEF and necessitate compensating adjustments to active tension. Zingaro et al. (2024, Journal of Computational Physics), using an alternative active tension model (RDQ20), similarly found it necessary to increase contractility parameter (a_XB) to achieve sufficient ejection, and explicitly report that no single parameter configuration simultaneously achieved physiological peak ejection rate and LVEF, a fundamental tension we also encountered. Together, these comparisons suggest that the elevated Tref in our model most likely reflects a combination of the absence of atrial filling, simplifications in pericardial constraint, and the lack of poroelastic behaviour, rather than volumetric locking alone. Additional investigations are needed as explained in the manuscript, and these factors are identified as open priorities for future development within our framework.

      We added a comparison with these three studies to the discussion section of the manuscript and have updated the limitations section to reflect this more nuanced account of the factors contributing to the elevated active tension scaling. Further work will be required to address these points.

      (b) Further, the authors employ the monodomain model for the simulation of the electrical excitation and relaxation on a relatively coarse grid with an approximate edge length of 1mm. This resolution is known to be insufficient for reliable results in organ-scale electrophysiology modeling.

      Our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached, as performed in Camps et al. (2024) (ref 33) using the tuneCV tool in monoAlg3D (https://github.com/rsachetto/MonoAlg3D_C/tree/master/scripts/tuneCV), which is similar to the tool in openCARP (described here: https://opencarp.org/documentation/examples/02_ep_tissue/03a_study_prep_tunecv).

      While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for simulations of ECGs in this study. We have noted this in the methods section.

      (3) Geometrical model and digital twin: The geometrical model, taken from a public cohort and calibrated to an ECG of another individual along with population-averaged values from a databank (UK Biobank), and unrelated measurements from surgical procedures, can hardly be considered a digital twin. Further, validation of the model was then performed against data from yet another cohort.

      We thank the reviewer for this point and welcome the opportunity to clarify our dataset choices. The use of multiple data sources was a deliberate methodological decision. While an ideal dataset for electromechanical model evaluation would combine full biventricular geometry, 12-lead ECG, invasive pressure measurements, and myocardial strain data from a single individual, no such dataset currently exists in the public domain, and acquiring it routinely would be impractical in clinical settings. Multi-source integration therefore reflects the realistic deployment scenario for future clinical translation of these tools.

      The specific choice of geometry was principled: the mesh associated with the ECG dataset that was available to us was truncated at the base due to the clinical acquisition protocol, which would have prevented physiologically realistic basal boundary conditions. The female Rodero geometry we chose provides full ventricular coverage and was selected on that basis.

      Demonstrating that a coherent, systematically evaluated framework can be constructed from compiled multi-modal data is itself a contribution because it makes the tools accessible to the wider community without requiring a single ideally acquired dataset.

      (4) Calibration procedure: There are apparent flaws in the calibration procedure, or it is not described in sufficient detail. The authors dedicate significant effort to motivating parameter ranges, but in the end they use mostly other parameters for the calibration process, aiming to maximize left ventricular ejection fraction. It is not clear whether the chosen parameters result in, e.g., physiological calcium traces or calibrated parameters that are within physiological ranges.

      Thank you for raising this point, which we have now clarified in the manuscript. The parameters that were chosen for the calibration process were based on the results of the sensitivity analyses.

      In addition, we have supplemented results Figure 1 with a subfigure F showing that the calcium transient and action potential durations fall within physiological ranges after calibration.

      (5) Goodness of fits, e.g., a direct comparison of the measured and the simulated ECG, are not provided to assess calibration quality.

      The calibrated model achieves QRS duration of 89 ms and QT interval of 360 ms, both of which fall within the healthy reference ranges compiled in Table 2, providing a biomarker-level assessment of calibration quality (Figure 1A). A full quantitative goodness of fit analysis of the simulated ECG morphology was performed following the methodology of Camps et al. [52], in which the same beat-averaged ECG was processed; we direct the reader to that work for full details rather than reproducing the analysis here.

      (6) Due to these limitations and weaknesses, the authors fall short of achieving some of their goals, particularly establishing credibility for the underlying computational framework and in reproducing healthy pressure-volume loops, and in achieving physiological simulations while using physiological or reported ranges for the calibrated parameters.

      For example, a key physiological requirement is that the right and left ventricular stroke volumes are approximately equal in a heart beating at a limit cycle, as the blood pumped by the right ventricle into the pulmonary circulation must match the amount pumped by the left ventricle into the systemic circulation. This balance is not achieved in this study.

      We thank the reviewer for identifying the stroke volume imbalance. We acknowledge the physiological requirement that, in a steady-state limit cycle, the right ventricular stroke volume must approximately equal the left ventricular stroke volume. However, since our model does not explicitly prescribe volumes, to achieve this, we would need to either explicitly tune active tension for the left and right ventricles separately, such as done in https://www.frontiersin.org/journals/physiology/articles/10.3389/fphys.2021.716597/full or develop a more sophisticated circulatory model and employ a multistep procedure that sequentially tunes circulatory dynamics, passive mechanics, and active contraction, such as done in https://www.biorxiv.org/content/10.64898/2025.12.11.693778v1.full. Both of which are beyond the scope of this paper.

      We note that, despite the absence of explicit RV calibration, the RV volumetric measures and pressures remain within physiological ranges, suggesting that the coupled biventricular mechanics are broadly plausible. As such, we have noted this limitation in our discussion section, and sign-posted to other studies where the stroke volume match is achieved.

      (7) The conclusive claim that "the study paves the way towards credible electromechanical cardiac Digital Twins" is not supported. The model exhibits non-physiological behavior, requires unsupported parameter alterations (such as a 10-fold active stress scaling), and does not represent a digital twin, as model data are drawn from various unrelated, non-patient-specific sources.

      We thank the reviewer for this comment, which gives us the opportunity to clarify our use of the term 'digital twin'. A cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning. This is a transformative goal for precision cardiology that the field is actively working towards, with credible, systematically validated electromechanical models as its essential foundation. To our knowledge, no published study in cardiac electromechanical modelling has simultaneously fulfilled all three requirements, and it is for this reason that the community often refers to the 'digital twin vision' rather than its realisation.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers in a single study. The model evaluation presented here is an example application of that framework. Importantly, the framework is not designed to certify a model as complete, but to provide a transparent audit of current capability by identifying where confidence is established and where further development is needed. In this sense, the limitations surfaced through this evaluation are themselves a contribution: they define open problems and priorities for the field.

      We have added a definition of the digital twin concept and the roadmap towards its realisation to the introduction and have updated the language throughout the manuscript to consistently reflect the distinction between the framework contribution and the model evaluation. We maintain that this transparent approach represents a meaningful step towards the digital twin vision.

      The specific limitations of the current model implementation are addressed in detail in the relevant sections of this response and in the updated manuscript, where we have substantially strengthened the verification and discussion components.

      Conclusion:

      Overall, this reviewer considers that the study requires a major revision, including improvements in numerical methods, modeling choices, and checks for physiological behavior. Nevertheless, the provided tables with averaged values from the UK Biobank and the presented validation strategy could be valuable to the research community.

      Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      We thank the reviewer for this comment.

      We have strengthened the quantitative aspect of the validation by reporting the peak simulated strain values for each component and comparing them against the physiological ranges compiled in Table 2. Specifically, the simulated peak strains were: E_ff ≈ -0.20, E_cc ≈ -0.15, E_rr ≈ +0.15, and E_ll ≈ -0.23. These show broad agreement with the in vivo reference ranges from Moulin et al. (2021), noting that the reference ranges are derived from a cohort of 30 subjects and therefore represent a relatively narrow population sample. Shortening strains (fibre and circumferential) are in good agreement, while radial strain is underestimated. We have noted this as a limitation. We have also indicated that a further validation would include a fully quantitative spatial comparison, such as a displacement error norm or voxel-wise strain difference map. This would require access to the raw image data and patient-specific geometry registration, which is beyond the scope of the current study.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      Thank you for raising this point. We have now included a section on model verification to the manuscript at where we perform benchmarking simulation using the Land (2015) passive inflation benchmark. We also provide a mesh subdivision analysis of the final calibrated model. Additional verifications and previous sensitivity analyses using the same numerical scheme with idealised ellipsoid geometries are also referenced in the verification section, to provide additionally confidence.

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Numerical analyses for the Alya solver used in this study has previously been published in works including Levrero et al (2021), which performed sensitivity analyses in a truncated ellipsoid geometry, and Santiago et al (2018), which demonstrated mesh convergence in a cantilever. We have updated the manuscript to point the reader to these studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major concerns:

      (1) Active stress scaling:

      The initial value for T_ref appears to be 120kPa * 10, which would be ten times the literature value fitted to human contraction data. Additionally, Table 2 lists a range of [1200-2400], which is 10 to 20 times the literature value.

      This discrepancy suggests that other model parameters, model assumptions, or the numerical scheme may be inadequate. In contrast, similar calibrations using comparable models (ToRORd-Land) in other works, such as Strocchi et al. [29], yielded T_ref values close to the literature value.

      We thank the reviewer for this comment. As discussed in our response to the public review comment 3a, the elevated T_ref scaling warrants explanation.

      We note that Strocchi et al. use a four-chamber geometry include atrial mechanics and a different pericardial constraint, any of which could contribute to differences in the required T_ref scaling. The elevated scaling in our model likely reflects a combination of factors including the absence of poro-elastic behaviour, simplifications in pericardial constraint, and the lack of atrial mechanics, rather than volumetric locking alone. We have added text to the discussion acknowledging this more explicitly and have flagged planned additional benchmarking of the dynamic orthotropic scheme as future work.

      (2) Non-physiological results, see Figure 1:

      In a healthy heart, RV stroke volume should approx. match LV stroke volume. This is clearly not the case in Figure 1B, where the RV EF is also notably low at 35%.

      Consequently, the study fails to reproduce healthy pressure-volume loops, undermining its claim to create a credible cardiac electromechanical digital twin. Hence, also the "Question of interest" posed in line 206 must be answered with a clear "No".

      Matching stroke volumes should be a primary calibration goal.

      We thank the reviewer for this comment. We agree that stroke volume balance is an important physiological criterion, and we have added it explicitly to the framework criteria in the updated manuscript, noting that our current model evaluation does not satisfy it. This is precisely the kind of transparent appraisal the V&V40 framework is designed to produce: a systematic accounting of which criteria are met and which require further development. A framework that only gets applied to models that pass all criteria would be selection-biased and less informative to the community.

      However, we respectfully disagree that the question of interest must be answered with a clear 'No'. We draw the reviewer's attention to the quantities of interest defined in the paper, which are predominantly left ventricular biomarkers, reflecting the intended scope of the calibration framework. The framework successfully reproduces these defined quantities of interest, and the LV pressure-volume loops, strain, volumes and ejection fraction are all within physiological ranges and well-matched to reference data. These quantities of interest were selected based on their clinical implications in cardiac diseases, as detailed in Table 3.

      We agree that stroke volume balance is an important physiological requirement for a fully calibrated biventricular model, and we have strengthened the future work and limitations section accordingly.

      (3) Inadequate numerical framework:

      (a) Monodomain model: The geometries from Rodero et al. [25] have an average edge length of 1mm. It is known that such a coarse resolution leads to inaccurate EP results. It is not mentioned if the authors refined that geometry to an appropriate resolution or used an Eikonal model to mitigate this issue.

      As explained earlier, our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface, as performed in Camps et al. (2024), and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached. We have added a figure in the appendix of this manuscript to show that by increasing the mesh resolution by one subdivision, we get virtually identical ECG simulations. While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for this study. We have noted this in the methods section.

      (b) Material law:

      - recent publications show that an unsplit deformation gradient for the anisotropic contribution is beneficial to reduce locking effects, see, e.g., Gueltekin et al. Computational Mechanics 63, no. 3 (2019): 443-53. https://doi.org/10.1007/s00466-018-1602-9.

      - K_ct is a penalty parameter to enforce some degree of incompressibility. Results are highly dependent on the grid size and the finite element formulation due to locking effects.

      As the authors write: "In our simulations, we saw that the LVEF was strongly sensitive to changes in the incompressibility of the tissue (Kct), such that an increase in compressibility of the myocardial tissue helped to increase LVEF." Which exactly points to the issue of locking effects.

      So an option would be to use a finer grid or a more adequate numerical scheme with quadratic finite elements, as eg. in [5] Fedele et al., or [6] Gerach et al,. or stabilized elements as in Karabelas et al. CMAME 394 (2022) https://doi.org/10.1016/j.cma.2022.114887.

      Overall, this does not point to "limitations in using ex vivo tissue measurements to represent in vivo function" but to limitations in the numerical setup. In fact, with an adequate numerical scheme, the simulations should be largely insensitive to the choice of this penalty parameter K_ct. See, e.g., Karabelas et al. above, where the authors varied K_ct from 650kPa to infinity (representing an incompressible material), and there is no visible influence on the PV loops.

      We investigated this point using the Alya solver, and we found that the mesh resolution did not alter the LVEF, and our benchmark simulations against Land (2015) did not show the existence of the volumetric locking issue that the reviewer refers to. It is possible, however, that such an effect exists in the elastodynamic orthotropic framework but not in the incompressible and transversely isotropic framework that the Land (2015) benchmarks were set up in. Future analyses could focus on performing additional benchmarking against more recent elastodynamic benchmarks, such as presented in Arostica (2025). We have updated the limitations text in our manuscript to reflect this and to cite relevant literature on this issue.

      (4) Boundary conditions:

      "This was a simplified version of the method [28], which uses an exponential decay formulation at the 'edge' of the pericardial constraint rather than a step function": I don't really see this in the cited work [28] which gives a spatially varying Robin-type boundary condition at the whole epicardium (i.e. regional scaling of normal springs stiffness based on image-derived motion from CT images) and not only at the edge.

      This is motivated by the fact that the pericardial tissue is in contact with various organs of different material properties. Not using spatially varying pericardial parameters is a limitation that might lead to non-physiological deformations, see also Pfaller et al. Biomechanics and Modeling in Mechanobiology 18 (2019): 503-29. https://doi.org/10.1007/s10237-018-1098-4.

      We thank the reviewer for this point and we have corrected the manuscript accordingly. To clarify: our implementation applies a uniform Robin spring constraint along the majority of the epicardial surface with zero constraint at the base, which is conceptually similar to Strocchi et al. [28]. The key difference is that Strocchi et al. use a smooth gradient transition from uniform constraint to zero constraint near the base, whereas our implementation uses an abrupt step transition. We acknowledge that a smooth spatially varying transition would more accurately represent the frictionless pericardial contact and have noted this as a limitation in the manuscript with reference to Pfaller et al. [41].

      Also check:

      - line 117: Gamma_valve_epi is introduced but not used. Was there any boundary condition defined on this valve plane?

      - the third equation, maybe (0,T] missing.

      - line 120: epicardium instead of endocardium.

      These errors have been corrected in the updated manuscript. No boundary conditions were applied on the epicardial surface of the valve plugs, the reference to gamma_valve_epi has been removed.

      (5) Reference geometry:

      The choice to scale the mesh to a lower volume for the unloading procedure seems questionable. This approach does not ensure that the reloaded mesh aligns with the mesh derived from image data. As a result, the geometry used for the simulations is no longer truly patient-specific.

      This mismatch is a significant limitation, as there are established methods available to achieve a proper unloaded configuration, as, e.g., in

      Marx et al. Journal of Computational Physics 463 (2022): 111266. https://doi.org/10.1016/j.jcp.2022.111266, and

      Regazzoni et al. Journal of Computational Physics 457 (2022): 111083. https://doi.org/10.1016/j.jcp.2022.111083.

      As our study aimed at creating a framework for calibration and validation in data-scarce scenarios such as it is often the case in the clinical context, using a compilation of multi-modal data from difference sources, rather than a specific method of personalisation, we did not feel it appropriate to invest significant energy to identify a patient-specific resting geometry, but rather felt that it was important for the resting geometry to fall within population values in terms of diastasis volume. We have clarified this issue in the manuscript and softened claims to Digital Twins in this study. The limitation has been addressed in the updated manuscript, and future work could further address this point.

      (6) Calibration procedure:

      There are apparent flaws in the calibration procedure, or it is not described in sufficient detail.

      We thank the reviewer for raising this point and we have substantially revised the calibration description in the manuscript to clarify the rationale behind each step.

      (a) Step 1: "Sample..." Why? kws and Cal50 are not the most significant parameters in the sensitivity analysis. Kct is a penalty parameter dependent on the numerical framework as described above; "ejection pressure threshold" was never mentioned, is it "P ejection LV" in Table 2? Aiming just for the highest LVEF might neglect non-physiological responses to parameter changes.

      While kws and Cal50 are not the single most significant parameters for LVEF in isolation, they were grouped in Step 1 because they affect both LVEF and peak systolic pressure simultaneously through cross-bridge cycling rate and residual active tension, making it necessary to sample them jointly rather than sequentially. Kct was included because myocardial compressibility affects wall thickening and therefore stroke volume. The ejection pressure threshold is P_ejection_LV in Table 1 and has now been described explicitly in the methods section. Regarding the concern about non-physiological responses: the action potential duration and active tension were monitored throughout calibration and verified to remain within physiological ranges, as now noted in the manuscript.

      (b) Step 2: As systolic pressure is directly dependent on arterial resistance for a 2-element Windkessel model, a uniform sampling approach might not be the best choice here.

      We acknowledge that uniform sampling may not be the most efficient approach for Step 2. However, since arterial resistance influences not only peak systolic pressure but also stroke volume and therefore LVEF, a more targeted approach focusing solely on pressure matching could compromise the LVEF achieved in previous steps. Uniform sampling allowed us to select the value that best balanced both quantities simultaneously.

      (c) Step 3: The authors mention in line 527: "A four-fold increase in GCaL caused an eight-fold increase in cellular active tension peak". An increase in active tension peak results in higher LVEF. So this step is likely to yield the upper boundary of the GCaL interval.

      The reviewer is correct that Step 3 tends to yield a high GCaL value. This was intentional — GCaL was used as a last resort to achieve physiological LVEF after Steps 1 and 2, since the model consistently undershot the target. The upper boundary of the sampled GCaL interval corresponds to a two-fold increase, which remains within the physiological variability bounds applied in previous studies. The resulting action potential duration was verified to remain within physiological ranges.

      (d) Step 4: Why again k_ws? It is not the most significant parameter in the SA.

      kws was resampled in Step 4 not to increase LVEF further, but to specifically target peak ejection rate and dP/dtmax, which were not adequately matched after Step 3. kws is the dominant parameter affecting these ejection dynamics biomarkers in the sensitivity analysis. Resampling at this stage allowed fine-tuning of ejection dynamics while maintaining the LVEF achieved in previous steps.

      (e) Step 5: As far as I can tell, the "diastolic volume change parameter" was mentioned the first time here.

      The diastolic volume change parameter C_pLAV has now been described in the methods section in the Phase 5 passive filling description, where it appears as the inverse of the penalty term controlling the rate of return to diastasis volume in the left ventricle.

      The whole calibration procedure seems to aim for the highest LVEF, and final values of the calibration parameters are not given.

      We note that the calibration procedure does not aim solely for the highest LVEF. As described above, the sequential strategy targets multiple quantities of interest in order of clinical importance: LVEF, peak systolic pressure, peak ejection rate, and peak filling rate, with each step designed to improve a specific subset of biomarkers without compromising those already matched. The final calibrated parameter values are reported in Figure 1F of the revised manuscript.

      (7) Novel features in this paper are actually scarce. A way more advanced calibration strategy with a whole heart model, emulators, and also the ToRORd-Land model was already presented in the study by Strocchi et al. [29]. The calibration to ECGs was presented by some of the same authors in Camps et al. [15], and the analysis of cellular effects was already published in several studies by the same group and in other publications, e.g., by the groups of Severi et al.

      The systematic compilation of credibility criteria spanning ECG morphology, pressure-volume characteristics, strain and displacement represents a novel contribution in itself, providing the field with a reusable evaluation framework. Furthermore, the present study is designed to yield mechanistic insight into how parameters at different scales influence both electrical and mechanical outputs simultaneously. This goal was not tackled in previous publications, which covered individual components, including ECG calibration in Camps et al. [15] and global sensitivity analysis with whole-heart models in Strocchi et al. [29] with no ECG consideration.

      Thus, the work by Camps et al. on ECG calibration was purely electrophysiological and did not investigate the influence of mechanical or haemodynamic parameters on ECG morphology in a fully coupled electromechanical framework. While the effect of mechanical parameters on ECG has been explored by others (e.g. Favino, 2016), this has not previously been examined alongside the relative importance of cellular, mechanical and haemodynamic parameters on pressure-volume characteristics within a single coupled framework. While Strocchi et al. present an emulation strategy, they did not address ECG biomarkers. This distinction is now stated explicitly in the introduction, where we position the present study relative to Camps et al. and Strocchi et al.

      (8) How could the calcium sensitivity Cal50 have such a drastic effect on diastolic function, i.e., filling and end-diastolic volume? As far as I understand from the description, the simulation starts with Phase 0 (loading), Phase 1 (atrial filling), and then in Phase 2, electrical activation ensues and active contraction develops, see also the section starting in line 165. Based on this description, I would expect the end-diastolic volumes to be identical across all Cal50 values. Or are the PV loops shown actually limit cycles established over simulations with multiple beats? This point wasn't explicitly clarified in the manuscript.

      Calcium sensitivity (Cal50) affects not only systolic active tension development but also diastolic residual active tension, i.e. the degree to which the muscle remains partially activated at end diastole. Higher Cal50 values increase this residual tone, effectively stiffening the myocardium during diastolic filling and reducing end-diastolic volume. This mechanism is well established as a contributor to diastolic dysfunction in heart failure [88]. We have clarified this in the manuscript and also clarified that the PV loops shown are single-beat simulations, not limit cycles, with the end-diastolic volume determined by the prescribed filling pressure alongside the passive and residual active stiffness of the myocardium.

      Minor concerns:

      (9) Line 29: The values provided: LVEF of 51%, EDV of 110 mL, and ESV of 50 mL are inconsistent. If these values are all related to the LV, the calculated LVEF should be approximately 54.55%, not 51%.

      The values quoted in the original abstract were rounded approximation, this has been corrected to report EDV=105 mL and ESV=51 mL, which are consistent with the simulated LVEF of 51%.

      (10) "Electromechanical cardiac Digital Twins have had broad applicability ..."

      Many of the cited works here are not true "Digital Twins" but rather static, non-patient-specific models of cardiac electromechanics. In some cases, the geometry may be derived from patient data, but this alone does not qualify the model as a digital twin.

      This sentence in the introduction has been rephrased as ‘Electromechanical cardiac models have had broad applicability...’. Furthermore, as stated earlier, we have removed explicit claims of Digital Twin from the paper while retaining the fact that this study provides a significant step towards rigorous credibility assessment of the high-fidelity electromechanical models that make Digital Twin construction possible.

      (11) While in the abstract and the conclusion, the authors mention "uncertainty quantification", it is mostly a sensitivity analysis that was performed in the paper.

      We have updated the text to say ‘sensitivity analysis’ where appropriate in the abstract, results, and conclusion, and replaced ‘uncertainty ranges’ with ‘variability ranges’ throughout. However, since the sensitivity analyses were performed over biologically informed ranges derived from population variability in the literature, the results are informative about how uncertainty in model inputs propagates to uncertainty in simulated biomarkers. We have therefore retained the framing of sensitivity analysis as a first step towards uncertainty quantification in the abstract and conclusion, and have added a clarifying sentence to the methods to this effect.

      (12) Line 98: As far as I can tell, the conduction velocity assigned to the endocardial surface - intended to mimic the Purkinje fiber network - is never specified. In the section beginning at line 294, only the transmural conduction velocities are reported.

      The endocardial conduction velocity has been specified in the methods section: Purkinje-myocardial junctions were modelled using a fast endocardial activation layer with isotropic conduction velocity of 300 cm/s.

      (13) Line 198, Table1:

      (a) "21/02/2025 11:09:00 AM" on two occasions is maybe not intended

      This has been removed.

      (b) For easing up comparisons, units should be consistent between the initial value and the literature ranges, e.g., PV control parameters, heart rate.

      Units have been made consistent between the initial values and literature ranges throughout Table 1.

      (14) Line 232: "... have already been used to calibrate and validation ...".

      This has been corrected.

      (15) Line 279, Table 2: This table of variability ranges is not entirely clear and could be improved:

      Table 2 has been combined with Table 1 such that the variability ranges sit next to the literature values, for ease of comparison.

      (a) "21/02/2025 11:09:00 AM" is maybe not intended.

      This has been removed.

      (b) use of units should be improved; sometimes it's given in the first column, sometimes in the second column (arterial resistance, compliance), then for k_epi it should be either kPa or kPa/cm.

      Units have been made consistent and the units for k_epi has been added in Table 1.

      (c) units should also be consistent throughout the paper, e.g. in Figure 1 E arterial resistance is Barye.ms/mL while in Table 2 it is mmHg.ms/mL.

      Barye has been removed and replaced by corresponding kPa values throughout the manuscript. This was in the original manuscript since the Alya simulation software were in units of cm, s, g, Barye.

      (d) it is also not clear how variability ranges were chosen; e.g., for arterial compliance,e literature ranges are 0.2-2.73 while the chosen range is [0.1,0.2].

      The previous ranges were chosen to achieve better LVEF. We have now updated the variability ranges to be purely based on literature values and updated the sensitivity analysis results. The ranges are now presented in Table 1 alongside the literature values for ease of comparison.

      (e) For Kct, the initial value in Table 1 is 5000kPa, the literature values are between 10 and 3333, and then the variability range is [10,500]? I guess there is a typo in one of these values.

      This has been corrected in the new Table 1.

      (f) Table 1 and 2 are in parts redundant.

      Table 1 and 2 have been combined into a single new Table 1.

      (15) Line 290: It should be uvc_l for the longitudinal coordinate.

      This has been corrected.

      (16) Line 388, Table 3, regarding values for pressure volume from reference [49]:

      (a) the number of participants is 800, including males and females; not only females, see also Table 12 https://jcmr-online.biomedcentral.com/articles/10.1186/s12968-017-0327-9/tables/12

      (b) why using female values here while having mixed sex for most of the others? Because the model is female?

      The reviewer is correct that reference [49] reports values from a mixed-sex cohort of approximately 800 participants. We used the female-specific values from Table 12 of that reference because the biventricular mesh used in this study was derived from a female subject, making sex-matched reference values the most appropriate comparison. This has been clarified in the manuscript.

      (17) Figure 5: What is Jup; why did you choose 0.93 x Jup as reference? Also in Figure 4, why did you use 0.93 x GCal as a reference?

      J_up refers to the SERCA<sup2+</sup> reuptake current, which has been relabelled as SERCA throughout the manuscript for consistency. The reference value of 0.93× was used because the sensitivity analysis sampled parameters uniformly between 50% and 200% of baseline using a fixed number of samples, and no sample fell exactly at 1.0×. The closest sampled value was 0.93×, which was therefore used as the reference. This has been clarified in the figure caption.

      (18) Line 377: The link to the GitHub repository does not work.

      This link has now been made publicly available.

      (19) Line 397, Table 3: for the sake of completeness, all abbreviations should be included: e.g., SVL, ESP, EDV, ESV are not included.

      This has been written out in full in the new Table 2.

      (20) Tick marks in many figures are not readable, e.g., Figure 5 and all the Figures in the appendix.

      Tick mark sizes and line widths have been increased across Figures 4, 5, and all appendix figures. The figures have been replotted and updated in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor Comments:

      (1) The provided GitHub link https://github.com/jennyhelyanwe/Alya_input_setup/ does not work, potentially because the repository is private. It would be nice to see the repository during the review.

      This link has now been made publicly available.

      (2) Table 3: Can you include the simulation outputs obtained for validation (with an error indication)? This would summarize the validation that's currently spread out over the results section.

      A new Table 3 has been added to the manuscript under the validation section, summarising the simulated values for all deformation and strain biomarkers alongside their reference ranges. The calibration and validate datasets are now reported separately in Tables 2 and 3, respectively.

      (3) Figure 1: Add axis labels to all plots.

      Axis labels have been added to all subplots in Figure 1 in the revised manuscript. Simulated pseudo-ECG amplitudes are normalised and therefore dimensionless.

      (4) Figure 2: Simulated and in vivo strains with exactly the same axes (size, range, ticks) and add grid lines to enable a comparison. Add the mean values of each in the other plot.

      The revised Figure 2 now includes the median in vivo strain values from Moulin et al. overlaid as a red dashed reference line on the simulation panels, enabling direct visual comparison. The simulated mean could not be overlaid on the in vivo panels as the original Moulin et al. figure data are not publicly available for replotting. Exact axis matching was not applied as this would cause some simulated curves to fall outside the visible range, obscuring the model behaviour.

      (5) Figure 3: The thickness (relative importance of the connections) is impossible to see in this plot. Instead of having gray background connections, remove them entirely below a certain threshold. Make the differences in thickness more pronounced or introduce a continuous color scale for the magnitude of the positive or negative correlation. Alternatively, you could rank the parameters from least to most important in each subfigure A-D and/or provide some numeric values.

      Figure 3 has been updated. All non-significant connections (|r| < 0.6 or p > 0.05) have been removed entirely, and gray lines have been removed in each subfigure, making the significant relationships clearer. A continuous blue-to-red colour scale has been applied to indicate the direction of correlation (blue: negative, red: positive), with line thickness proportional to the magnitude of the r-value.

      (6) Figure 3 and Table 2: Why were material parameters b, bf, bs, and bfs omitted from this study (but included a, af, as, and afs)?

      The b parameters (b, bf, bs, bfs) appear in the exponent of the Holzapfel-Ogden constitutive law and are strongly coupled to the a parameters (a, af, as, afs), which carry units of kPa. In practice, the b parameters can only be reliably identified from ex vivo multiaxial stretch experiments, whereas the a parameters can be estimated from clinical imaging data. Since our study focuses on calibration and validation in a clinical data setting, we included only the a parameters in the sensitivity analysis, consistent with previous personalisation studies.

      (7) Figure 4: What do the dotted lines represent?

      The dotted lines in Figure 4E highlight the increased longitudinal shortening with increasing GCaL, showing the basal plane moving towards the apex while the apical position remains unchanged due to the pericardial constraint. This has been clarified in the figure caption.

      (8) Figures 4, 5, A2-45: Can you use a continuous color scale (e.g., from blue to red) for low to high parameter uncertainty?

      A continuous blue-to-red colour scale has been applied to Figures 4, 5, and all appendix figures A2–A6, where blue indicates the lowest parameter value and red indicates the highest. A colour bar has been added to each figure for reference.

    1. eLife Assessment

      The study presents important findings from a very rich EEG-fMRI dataset including 107 participants, which was collected during nocturnal naps. Using overall solid methods, the authors link activity in memory related brain regions (e.g., hippocampus, thalamus, and medial prefrontal cortex), as well as their functional connectivity, to the occurrence of canonical sleep rhythms (spindles and slow oscillations) during non-rapid eye movement (NREM) sleep. This work will be of broad interest to researchers studying sleep, memory, and related domains.

    2. Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Comments on latest version:

      The authors have substantially revised their manuscript and sufficiently addressed all of my previous concerns. I have no further comments.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Thank you for this encouraging assessment. We appreciate your recognition of the value of the dataset and of the questions it allows us to address. Below, we respond to each of your points directly and revise the manuscript accordingly.

      Comments on revisions:

      Re 1: The revised introduction now cites a couple of papers but discusses them only very superficially, lumping together several studies with very different key results. This is still not very informative for the reader and does not sufficiently acknowledge previously published work. Here are two examples to illustrate this:

      (a) "These studies have generally reported that slow oscillations are associated with widespread cortical and subcortical BOLD changes, whereas spindles elicit activation in the thalamus, as well as in several cortical and paralimbic regions." Several studies even showed e.g., a clear activation of the hippocampus and parahippocampal gyrus associated with spindles, not just the thalamus

      Thank you for this comment. We agree that our previous sentence was too broad and did not sufficiently reflect the range of findings in the sleep literature. We have therefore rewritten the Introduction to state explicitly that spindle-related BOLD changes have been reported not only in the thalamus, but also in cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).

      Introduction, Page 3-4, Lines 58-62

      “Consistent with this view, prior human EEG-fMRI studies have reported spindle-related activation not only in the thalamus, but also in the hippocampus and adjacent parahippocampal gyrus (Bergmann et al., 2012; Schabus et al., 2007). Spindle-related activity has also been linked to striatal engagement, suggesting a broader network that may support memory-related processing during sleep (Fogel et al., 2017).”

      Introduction, Page 4, Lines 71-78

      “Previous EEG-fMRI studies on sleep have examined both global sleep characteristics (Hale et al., 2016; Moehlman et al., 2019) and the neural correlates of specific waves, including slow oscillations and spindles. These studies have generally shown that slow oscillations are associated with widespread cortical and subcortical BOLD changes (Czisch et al., 2009; Ilhan-Bayrakcı et al., 2022; Picchioni et al., 2011), whereas spindles have been linked not only to thalamic activation but also to cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      Introduction, Page 5, Lines 103-106

      “This coupling was associated with increased activation in both the thalamus and hippocampus, with functional connectivity patterns suggesting thalamic coordination of hippocampal-cortical communication, in line with prior EEG-fMRI studies of spindle-related activity (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      (b) "Although these findings provide valuable insights into the BOLD correlates of sleep rhythms, they often do not employ sophisticated temporal modeling (Huang et al., 2024) [, ...]." - previous studies have used e.g., spindle event-related regressors with individual spindle amplitudes as parametric modulators, first and second order derivatives of the HRF function, as well as PPI connectivity analyses, which I would consider rather sophisticated temporal modelling.

      We agree that several previous studies have already employed sophisticated modelling approaches, including parametric modulation, HRF derivatives, and PPI analyses (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Our intention was not to suggest that such methods are absent from the literature.

      Rather, we aimed to highlight that most prior work has focused on modelling individual SO or spindle events, whereas explicit modelling of their temporal interaction (e.g., SO-spindle coupling) has been less commonly addressed. We have revised the sentence to clarify this point more precisely.

      Introduction, Page 4 Lines 78-82

      “Although these findings provide important insight into the BOLD correlates of sleep rhythms, most previous studies have focused on individual oscillatory events rather than explicitly modelling their temporal interaction (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Only a few recent studies have begun to examine coupling between rhythms directly, for example Huang et al. (2024).”

      Re 4+9: The short overall recordings in some subjects on the one hand and the large number of spindles and SOs detected in N1 sleep stages are still highly concerning, in fact even more so, now that the actual numbers have been provided in the Supplementary Tables. Either the sleep staging or the detection of SO and spindle events must be incorrect. I understand that for specific EEG analysis and fMRI modelling purposes sometimes slightly different thresholds are used as compared to clinical sleep staging, but several parameters here are alarmingly off.

      (a) Given that proper NREM sleep (N2+N3) is the relevant stage for the analyses conducted in this paper, some of the N2+N3 durations are very short (eg 7-8 min) while those subjects' results have the same impact on the group level analyses as those with >100 min of N2+N3. Either subjects with very little relevant data (not overall recording time but N2+N3 time) should be excluded or weighting subject data for the group analyses according to the amount od contributed data should be done.

      Thank you for the suggestion. It is true that participants with very little N2/3 sleep could contribute noisier subject-level estimates to the group analysis. We therefore checked this directly. Only three participants contributed less than 10 min of N2/3 sleep, and excluding them did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant after exclusion, t<sub>(103)</sub> =2.50, p = 0.0071, compared with t<sub>(106)</sub> = 2.50, p = 0.0070 in the full sample. We have added this control analysis to the Results so that the robustness of the group findings is explicit in the manuscript.

      Results, Page 11-12, Lines 238-250

      To ensure the results were not driven by individual differences or parameter selection, we conducted a series of control analyses. First, we excluded participants with less than 10 minutes of N2/3 sleep. Only three participants met this criterion, and their exclusion did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant (t<sub>(103)</sub> = 2.50, p = 0.0071), comparable to the full sample (t<sub>(106)</sub> = 2.50, p = 0.0070). Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st-80th percentile; Fig. S6). Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).

      (b) The authors argue that the SO and spindle detection algorithms are valid since widely used and that they were developed for N2+N3 stages, which is why they will also detect events in other stages: "While, because the detection methods for SO and spindle are based on percentiles, this method will always detect a certain number of events when used for other stages (N1 and REM) sleep data, but the differences between these events and those detected in stage N23 remain unclear." I do agree that with very liberal thresholds, also SO and spindle vents may be detected in other stages, but it shouldn't be that many. If the percentiles of amplitude thresholds were defined based on properly scored N2+N3 stages only, very few events should be detected (erroneously!) in N1, as the occurrence of K-complexes (isolated SOs) and spindles per definition makes it N2, and during REM sleep only very few spindles and SOs are allowed to occur, without scoring it NREM instead. For the first subject (just as example, but with similar numbers for the rest of the sample), reveals as many as 60 SOs and 31 spindles within 8 min of N1 sleep (Table S2) as well as 13 SOs and 7 spindles within 2 min of REM sleep (Table S4). These numbers are completely unrealistic and question the correctness of the sleep staging as well as the physiological relevance of the EEG graphoelements identified as SO and spindles. It also completely undermines the interpretability of the respective event regressors for the fMRI analyses.

      (c) Likely, given the large numbers of coupled SO-spindle events and the apparently very low amplitude criteria for event identification, also the number of SO-spindle couplings is likely severely overestimated.

      We thank the reviewer for raising this important point. We agree with you that the original stage-wise percentile thresholding could inflate the apparent number of SOs and spindles outside N2/3 sleep. In the original analysis, the thresholds were estimated separately within each sleep stage. As you point out, this procedure can force the detector to label a relatively large number of events in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 SOs or spindles. We have therefore revised the detection procedure. Following your concern and Reviewer 2’s suggestion, the SO and spindle thresholds are now defined only from N2/3 sleep within each participant, where SOs and spindles are most abundant and physiologically expected to occur. These fixed N2/3-derived thresholds were then applied unchanged to N1 and REM for descriptive reporting. This avoids the artificial normalisation of event detection across sleep stages that can arise when each stage has its own percentile threshold. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We agree with you that the remaining detections in N1 and REM should not be treated as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We therefore report them only as descriptive detector outputs obtained under a fixed N2/3-derived threshold (see Table S2, S4 in the revised manuscript). We do not use them to support any physiological claim about SO-spindle coupling in N1 or REM.

      This point is also important for the fMRI analyses. You are right that inflated N1 or REM detections would undermine the interpretability of event regressors if those detections entered the EEG-informed fMRI models. They did not. All EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep, where SOs, spindles, and their coupling are physiologically expected and where the detection thresholds were defined. Thus, the central fMRI event regressors were based only on N2/3 events, not on detections from N1 or REM.

      We also agree with you that the absolute number of detected SO-spindle couplings depends on the chosen detection threshold. For this reason, we tested whether the main EEG-fMRI result depended on the specific detector setting. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We therefore do not argue that the detector provides a uniquely correct absolute count of SOs, spindles, or coupling events in every sleep stage. Our conclusion is more specific. The main N2/3 EEG-fMRI finding is robust across a reasonable range of SO detection thresholds, detections in N1 and REM are reported only descriptively, and the physiological interpretation of SO-spindle coupling is restricted to N2/3 sleep.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Re 10: The rationale for using a lateralized frontal electrode (F3) for both SO (should have been at least bilateral or central) and spindle detection (should have been a centro-parietal electrode) is not convincing. Other EEG-fMRI spindle or SO papers have used a number of frontal (SO) or centro-parietal (spindles) electrodes averaged or even approaches including all EEG electrodes. Searching events with low thresholds at suboptimal recording sites does not dot this highly valuable dataset justice.

      We thank the reviewer for this important comment. We agree that this choice is more sensitive to frontal SOs than to the centro-parietal fast spindle component. Our choice of F3 was driven by the practical constraints of prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to FCz can have reduced signal contrast relative to the reference, and electrodes near the vertex are also more vulnerable to prolonged pressure against the MRI head coil when participants sleep supine for several hours. In this setting, frontal electrodes provided more stable signal quality across the recording. Because the EEG events were used primarily as temporal markers for fMRI modelling, our priority was to obtain reliable event timing during N2/3 sleep rather than to estimate the full scalp topography of SOs and spindles.

      We also agree with your concern that this valuable dataset would ideally be analysed with multichannel detection strategies. To test whether the main result depended on the single lateralised F3 site, we repeated the main EEG-informed fMRI analysis using Fz, a midline frontal electrode. The result was unchanged. Hippocampal activation during SO-spindle coupling remained significant when events were detected from Fz, t<sub>(106)</sub> = 2.47, p = 0.0076, closely matching the original F3-based result, t<sub>(106)</sub> = 2.50, p = 0.0070. This control analysis does not remove the limitation that centro-parietal fast spindles may be underrepresented, and we do not claim that it does. It does show, however, that the main hippocampal fMRI finding is not driven by idiosyncratic detections from one lateralised frontal electrode. We have made this clearer in the revised manuscript.

      Finally, your concern about low thresholds is also important. As described in our response above, the revised analysis now defines SO and spindle thresholds from N2/3 sleep and applies these thresholds uniformly for descriptive comparisons across stages. We also tested the robustness of the hippocampal fMRI result across SO detection thresholds, and the effect remained significant across the 71st to 80th percentile range. We have therefore narrowed the interpretation in the revised manuscript. The main EEG-fMRI result reflects BOLD activity associated with frontal-channel-detected SO-spindle coupling during N2/3 sleep, rather than a full multichannel characterisation of all SO and spindle topographies.

      Results, Page 12, Lines 247-250

      “Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).”

      Discussion, Page 18, Lines 380-394

      “Third, sleep oscillation detection was based on a single frontal electrode. This choice improved signal stability and event timing in the prolonged simultaneous EEG-fMRI setting, but it did not exploit the full multichannel EEG information and cannot characterise the full spatial distribution of SOs and spindles. In particular, F3-based detection may be more sensitive to frontal SOs and frontal sigma activity than to the centro-parietal fast spindle component. We therefore interpret the EEG-informed fMRI results as reflecting BOLD activity associated with frontal-channel-detected SOs, spindles, and their coupling during N2/3 sleep. Future studies using multichannel or source-informed detection strategies, with separate treatment of slow and fast spindles, will be better suited to capture the spatial dynamics of these sleep oscillations. Fourth, the use of large anatomical ROIs may mask subregional contributions of specific thalamic nuclei or hippocampal subfields. Finally, without a memory task, we cannot establish a direct behavioral link between sleep-rhythm-locked activation and memory consolidation. Future studies combining ultra-high-field fMRI or iEEG with cognitive tasks, as well as multichannel or source-informed detection strategies that separately characterize slow and fast spindles, will be better suited to refine our understanding of subregional network dynamics and the functional significance of sleep oscillations.”

      Methods, Page 25, Lines 556-566

      “It is worth noting that the primary aim of EEG rhythm detection was to identify reliable event times for EEG-informed fMRI modelling. Detection was performed on the F3 electrode because this channel provided stable signal quality during prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to this reference, and electrodes near the vertex that were in prolonged contact with the head coil during supine sleep, were more susceptible to reduced signal contrast, impedance drift, and pressure-related degradation of electrode-scalp contact. We therefore used F3 as a pragmatic choice to maximize reliable event timing in N2/3 sleep. This choice was not intended to characterise the full scalp topography of SOs or spindles, and it may underrepresent the centro-parietal fast spindle component. As a sensitivity analysis, we repeated the main EEG-informed fMRI GLM using Fz, a midline frontal electrode, with the same detection and modelling procedure.”

      Re 7: It is not clear to me why/how larger voxels would reduce susceptibility-related distortions and partial volume effects. Usually, the opposite is true. This should be elaborated.

      What we meant was that we chose a relatively large voxel size to preserve signal-to-noise ratio and whole-brain coverage within a feasible repetition time for a long overnight EEG-fMRI protocol. This choice is useful for maintaining BOLD sensitivity in sleep recordings, where head motion, physiological noise, and participant comfort are major practical constraints. We agree that it may not be accurate to describe it as reducing susceptibility-related distortion or partial volume effects.

      We have rewritten the Methods to state this trade-off directly. The voxel size of 3.5 × 3.5 × 4.2 mm<sup>3</sup> allowed whole-brain coverage with a TR of 2000 ms, which was important for modelling sleep-rhythm-related BOLD responses across the whole brain during prolonged nocturnal recordings. A smaller voxel size would have improved spatial specificity, but would also have required either a longer TR, reduced brain coverage, or lower SNR, none of which would have been ideal for the present EEG-fMRI sleep design. We now explicitly acknowledge the cost of this choice.

      Methods, Page 21 Lines 453-463

      “For the functional scans, whole-brain images were acquired using a T2*-weighted gradient echo-planar imaging (EPI) sequence sensitive to the BOLD contrast. The sequence parameters were as follows: 33 slices in interleaved ascending order, TR = 2000 ms, TE = 30 ms, voxel size = 3.5 × 3.5 × 4.2 mm<sup>3</sup>, FA = 90°, matrix = 64 × 64, gap = 0.7 mm. A relatively large voxel size was chosen to preserve signal-to-noise ratio while maintaining whole-brain coverage within a feasible repetition time. This compromise was important for the prolonged overnight EEG-fMRI sleep protocol, where head motion, physiological noise, participant comfort, and sustained acquisition stability are substantial practical constraints (Bodurka et al., 2007; Laufs et al., 2008). A smaller voxel size would have improved spatial specificity, but would have required either a longer repetition time, reduced brain coverage, or lower signal-to-noise ratio.”

      Reviewer #2 (Public review):

      In this study, Wang and colleagues aimed to explore brain-wide activation patterns associated with NREM sleep oscillations, including slow oscillations (SOs), spindles, and SO-spindle coupling events. Their findings reveal that SO-spindle events corresponded with increased activation in both the thalamus and hippocampus. Additionally, they observed that SO-spindle coupling was linked to heightened functional connectivity from the hippocampus to the thalamus, and from the thalamus to the medial prefrontal cortex-three key regions involved in memory consolidation and episodic memory processes.

      This study's findings are timely and highly relevant to the field. The authors' extensive data collection, involving 107 participants sleeping in an fMRI while undergoing simultaneous EEG recording, deserves special recognition. If shared, this unique dataset could lead to further valuable insights.

      Thank you for this encouraging assessment. We appreciate your recognition of the effort involved in collecting this simultaneous EEG-fMRI sleep dataset. Below, we respond directly to your remaining concern.

      Comments on revisions:

      The authors' efforts in revising the manuscript and addressing the reviewers' comments are certainly commendable. However, I remain concerned about potential issues in detecting sleep-related oscillations (SOs, spindles, and consequently coupled SO-spindle events), which may arise due to suboptimal parameter selection or inaccurate sleep staging, potentially impacting all subsequent analyses.

      A review of Supplementary Tables 1-4 reveals an unusually high number of detected SOs and spindles during sleep stage N1 and REM sleep. While the authors correctly note that a percentile-based detection approach will always identify a certain number of events across sleep stages, the particularly high counts in N1 and REM are concerning. To mitigate the limitations of this method, the authors could have performed event detection independently of sleep stages (i.e., across the entire dataset for each participant) and subsequently assigned the detected events to the corresponding sleep stages. If the event counts in N1 and REM remained disproportionately high, this would indicate a fundamental issue with the detection procedure.

      In the previous version, thresholds were estimated separately within each sleep stage. As you point out, this can force the detector to identify a relatively large number of SOs and spindles in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 events.

      We have therefore revised the detection procedure so that event detection is no longer based on separate stage-wise thresholds. Following the logic of your suggestion, we first defined a fixed threshold for each participant and then assigned the detected events to their corresponding sleep stages afterwards. We used N2/3 sleep to define the SO and spindle thresholds because this is the stage in which these events are physiologically expected and most reliably observed. These same N2/3-derived thresholds were then applied unchanged to N1 and REM. This avoids the circularity of forcing a percentile-defined number of detections within each sleep stage.

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We also agree with you that the remaining detections in N1 and REM should not be interpreted as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We now state this explicitly in the manuscript. They are reported only as descriptive detector outputs under the fixed N2/3-derived criterion. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      We also would like to clarify our sleep staging procedure. The sleep staging was first performed using an established automated algorithm, YASA toolkit (Vallat & Walker, 2021), and then manually reviewed by two sleep experts. More importantly for the central results, all EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep. Thus, N1 and REM detections did not enter the event regressors used for the main fMRI analyses and do not affect the interpretation of the hippocampal or thalamic findings.

      Finally, we agree that the absolute number of SO-spindle coupling events depends on the detection threshold. We therefore tested whether the main fMRI result depended on the specific SO threshold. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We have made this clearer in the revised manuscript.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Reviewer #3 (Public review):

      Summary:

      Wang et al., examined the brain activity patterns during sleep, especially when locked to those canonical sleep rhythms such as SO, spindle, and their coupling. Analyzing data from a large sample, the authors found significant coupling between spindles and SOs, particularly during the up-state of the SO. Moreover, the authors examined the patterns of whole-brain activity locked to these sleep rhythms. The authors next investigated the functional connectivity analyses, and found enhanced connectivity between the hippocampus and the thalamus and the medial PFC. These results reinforced the theoretical model of sleep-dependent memory consolidation, such that SO-spindle coupling is conducive for systems-level memory reactivation and consolidation.

      Strengths:

      There are obvious strengths in this work, including the large sample size, state-of-the-art neuroimaging and neural oscillation analyses, and the richness of results. The results now inform hemodynamic neural activity that coincided with SO-spindle couplings.

      Weaknesses:

      My earlier comments were about the inability to make inferences on memory given the lack of memory tasks, and the weakness in using the open-ended cognitive state decoding.

      Comments on revisions:

      The current revision has addressed these major concerns. The authors expanded discussions regarding the theoretical implications of the work in a more nuanced manner.

      Thank you for taking the time to re-evaluate the manuscript. We are pleased that the revised Discussion now reads as more nuanced, especially in relation to the limits of the memory-related interpretation. Your earlier comments helped us sharpen both the claims and the framing, and we are grateful for that.

      References:

      Bergmann, T. O., Mölle, M., Diedrichs, J., Born, J., & Siebner, H. R. (2012). Sleep spindle-related reactivation of category-specific cortical regions after learning face-scene associations. Neuroimage, 59(3), 2733-2742.

      Bodurka, J., Ye, F., Petridou, N., Murphy, K., & Bandettini, P. A. (2007). Mapping the MRI voxel volume in which thermal noise matches physiological noise—implications for fMRI. Neuroimage, 34(2), 542-549.

      Caporro, M., Haneef, Z., Yeh, H. J., Lenartowicz, A., Buttinelli, C., Parvizi, J., & Stern, J. M. (2012). Functional MRI of sleep spindles and K-complexes. Clinical neurophysiology, 123(2), 303-309.

      Czisch, M., Wehrle, R., Stiegler, A., Peters, H., Andrade, K., Holsboer, F., & Sämann, P. G. (2009). Acoustic oddball during NREM sleep: a combined EEG/fMRI study. PloS one, 4(8), e6749.

      Fogel, S., Albouy, G., King, B. R., Lungu, O., Vien, C., Bore, A., Pinsard, B., Benali, H., Carrier, J., & Doyon, J. (2017). Reactivation or transformation? Motor memory consolidation associated with cerebral activation time-locked to sleep spindles. PloS one, 12(4), e0174755.

      Hahn, M. A., Heib, D., Schabus, M., Hoedlmoser, K., & Helfrich, R. F. (2020). Slow oscillation-spindle coupling predicts enhanced memory formation from childhood to adolescence. Elife, 9, e53730.

      Hale, J. R., White, T. P., Mayhew, S. D., Wilson, R. S., Rollings, D. T., Khalsa, S., Arvanitis, T. N., & Bagshaw, A. P. (2016). Altered thalamocortical and intra-thalamic functional connectivity during light sleep compared with wake. Neuroimage, 125, 657-667.

      Helfrich, R. F., Lendner, J. D., Mander, B. A., Guillen, H., Paff, M., Mnatsakanyan, L., Vadera, S., Walker, M. P., Lin, J. J., & Knight, R. T. (2019). Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans. Nature Communications, 10(1), 3572.

      Helfrich, R. F., Mander, B. A., Jagust, W. J., Knight, R. T., & Walker, M. P. (2018). Old brains come uncoupled in sleep: slow wave-spindle synchrony, brain atrophy, and forgetting. Neuron, 97(1), 221-230. e224.

      Huang, Q., Xiao, Z., Yu, Q., Luo, Y., Xu, J., Qu, Y., Dolan, R., Behrens, T., & Liu, Y. (2024). Replay-triggered brain-wide activation in humans. Nature Communications, 15(1), 7185.

      Ilhan-Bayrakcı, M., Cabral-Calderin, Y., Bergmann, T. O., Tüscher, O., & Stroh, A. (2022). Individual slow wave events give rise to macroscopic fMRI signatures and drive the strength of the BOLD signal in human resting-state EEG-fMRI recordings. Cerebral Cortex, 32(21), 4782-4796.

      Laufs, H., Daunizeau, J., Carmichael, D. W., & Kleinschmidt, A. (2008). Recent advances in recording electrophysiological data simultaneously with magnetic resonance imaging. Neuroimage, 40(2), 515-528.

      Maingret, N., Girardeau, G., Todorova, R., Goutierre, M., & Zugaro, M. (2016). Hippocampo-cortical coupling mediates memory consolidation during sleep. Nature Neuroscience, 19(7), 959-964.

      Moehlman, T. M., de Zwart, J. A., Chappel-Farley, M. G., Liu, X., McClain, I. B., Chang, C., Mandelkow, H., Özbay, P. S., Johnson, N. L., & Bieber, R. E. (2019). All-night functional magnetic resonance imaging sleep studies. Journal of neuroscience methods, 316, 83-98.

      Ngo, H.-V., Fell, J., & Staresina, B. (2020). Sleep spindles mediate hippocampal-neocortical coupling during long-duration ripples. Elife, 9, e57011.

      Ngo, H. V., Martinetz, T., Born, J., & Molle, M. (2013). Auditory closed-loop stimulation of the sleep slow oscillation enhances memory. Neuron, 78(3), 545-553.

      Picchioni, D., Horovitz, S. G., Fukunaga, M., Carr, W. S., Meltzer, J. A., Balkin, T. J., Duyn, J. H., & Braun, A. R. (2011). Infraslow EEG oscillations organize large-scale cortical–subcortical interactions during sleep: a combined EEG/fMRI study. Brain research, 1374, 63-72.

      Schabus, M., Dang-Vu, T. T., Albouy, G., Balteau, E., Boly, M., Carrier, J., Darsaud, A., Degueldre, C., Desseilles, M., & Gais, S. (2007). Hemodynamic cerebral correlates of sleep spindles during human non-rapid eye movement sleep. Proceedings of the National Academy of Sciences, 104(32), 13164-13169.

      Schreiner, T., Kaufmann, E., Noachtar, S., Mehrkens, J.-H., & Staudigl, T. (2022). The human thalamus orchestrates neocortical oscillations during NREM sleep. Nature Communications, 13(1), 5231.

      Schreiner, T., Petzka, M., Staudigl, T., & Staresina, B. P. (2021). Endogenous memory reactivation during sleep in humans is clocked by slow oscillation-spindle complexes. Nature Communications, 12(1), 3112.

      Staresina, B. P., Bergmann, T. O., Bonnefond, M., van der Meij, R., Jensen, O., Deuker, L., Elger, C. E., Axmacher, N., & Fell, J. (2015). Hierarchical nesting of slow oscillations, spindles and ripples in the human hippocampus during sleep. Nature Neuroscience, 18(11), 1679-1686.

      Staresina, B. P., Niediek, J., Borger, V., Surges, R., & Mormann, F. (2023). How coupled slow oscillations, spindles and ripples coordinate neuronal processing and communication during human sleep. Nature Neuroscience, 1-9.

      Vallat, R., & Walker, M. P. (2021). An open-source, high-performance tool for automated sleep staging. Elife, 10.

    1. eLife Assessment

      By mining the Logan assemblage of the Sequence Read Archive to reveal substantial papillomavirus diversity, this valuable study establishes a framework for integrating viral discovery with host, geographic, and ecological metadata. The evidence for identifying novel papillomavirus sequences is convincing, supported by large-scale sequence searches, established L1-based typing criteria, phylogenetic analyses, protein annotation, and structural comparisons. The broader host-association and ecological conclusions are more tenuous due to heterogeneous public metadata and uneven sampling, and would benefit from clearer discussion of potential biases, contamination, endogenous viral elements, and alternative explanations. The work should be of interest to a broad range of colleagues in the areas of microbiome research, virology, and viral evolution.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript titled "Petabase-scale papillomavirus discovery" capitalizes on Logan assemblages to identify novel and known PVs.

      This is a brilliant use of Logan assemblages, and this paper highlights the use of this, plus also shows that one person's trash is another one's gold. Superb paper and I applaud the authors for starting with the PVs as easier to identify due to their set of genes coupled with the conserved L1 protein and associated typing for PVs (10% pairwise identity threshold for identification of new PV types).

      Strengths:

      This study highlights the hidden gems in public resources, especially if mined properly. Thanks to Logan assemblages, this is possible, and this manuscript highlights this with their data mining of papillomaviruses, identifying known and novel PVs in pangolins, lizards, fish, and white rhinos.

      Weaknesses:

      None identified

    3. Reviewer #2 (Public review):

      This study applies the Logan assemblage of the Sequence Read Archive to investigate papillomavirus diversity at a large scale. The authors combine sequence similarity searches, phylogenetic analyses, protein annotation, structural comparisons, and metadata integration to identify and characterize novel papillomavirus sequences. Beyond papillomavirus sequence discovery, the study develops a framework for associating viral sequences with host, geographic, and ecological metadata through the integration of both established and recently developed large-scale public sequence repositories. This framework may provide a foundation for similar investigations across other viral taxa.

      An important aspect of the work is the systematic processing and integration of metadata across a large number of sequencing libraries. The approaches developed to aggregate, curate, and interpret host and environmental information may be broadly applicable to future large-scale studies of other viral families. At the same time, interpretations based on metadata-derived ecological and host-association patterns should be considered in the context of the inherent limitations of public sequencing repositories, including uneven taxonomic representation, heterogeneous sampling strategies, laboratory-derived samples, and variable sequencing protocols. The comments below primarily address methodological details, interpretation of ecological patterns, and opportunities to further strengthen the robustness of the analyses and conclusions.

      Pages 136-137: The manuscript focuses primarily on L1-containing contigs. Could the authors provide a summary of libraries containing other PV hallmark genes (e.g., E1 or E2) but lacking full-length L1 sequences? This would help assess how much additional PV diversity may remain inaccessible under an L1-centered framework. Furthermore, the manuscript does not provide a detailed discussion of technical or biological explanations for libraries with detected PV sequences but lacking L1 sequences. Could such cases represent incomplete assemblies, low-abundance infections, highly divergent PVs, or endogenous papillomavirus-derived elements? Clarifying these possibilities would help readers interpret the biological significance of these detections.

      Page 141: The rationale for performing the search in two sequential Logan releases is unclear. Does Logan v1.1 fully supersede v1.0? If so, why was the initial search performed on v1.0 rather than directly on v1.1? Providing a clearer description of the differences between the two database versions and explaining how these differences motivated the two-round search strategy would improve reproducibility.

      The data could have been explored at the libraries' read level. The analysis is currently limited to presence/absence and diversity patterns derived from assembled PV contigs. However, the identified PV-positive libraries provide an opportunity to explore abundance-related metrics. Read counts, coverage estimates, or other measures of sequence representation could be used to characterize PV abundance within libraries, providing additional ecological context and helping distinguish low-level incidental detections from strongly represented infections. In addition, the curated PV sequence dataset generated in this study could serve as a reference for targeted read-mapping analyses. Aligning reads from a subset of libraries classified as PV-negative may help determine whether PV sequences are present below assembly detection thresholds. Such an analysis could provide valuable insights into the sensitivity of assembly-based virus discovery approaches and help establish practical coverage or read-count thresholds for detecting low-abundance papillomaviruses. These results could have important implications for future surveillance, clinical, and environmental studies aimed at PV detection.

      Pages 155-156: Clustering was performed using sequence identity, while host, geographic, and ecological annotations were assigned from representative centroids. How frequently did clusters contain sequences associated with conflicting metadata (e.g., distinct hosts or geographic regions), and how were such cases handled?

      Page 160: The authors state that the 70% query coverage threshold was selected based on an observed bimodal distribution. Could this analysis be shown explicitly (e.g., in a supplementary figure), and could the authors discuss how sensitive the number of novel PV calls could have been to alternative coverage thresholds?

      Pages 167-168: The rationale for clustering highly divergent sequences at 60% nucleotide identity should be explained. Is this threshold associated with established genus-level classifications in this viral family, or was it chosen empirically?

      Pages 167-168: For the 45 sequences lacking nucleotide-level matches, did the authors investigate amino-acid similarity to known PV L1 proteins? Such analyses would help determine whether these sequences represent deeply divergent PVs or potentially more distant viral lineages.

      Pages 181-183: The biological interpretation of host-associated PV diversity may depend on library type. Could the authors summarize the proportion of samples originating from field collections, laboratory animals, cell culture systems, or experimental infections?

      Pages 243-248: Could the observed geographic and ecological patterns be influenced by laboratory-derived samples? Distinguishing field-collected samples from laboratory, captive, or experimental material would strengthen the ecological interpretations.

      Pages 254-261: The biome analyses focus on PV occurrence. An analysis of host composition across biomes would be highly informative and could help disentangle whether observed patterns reflect PV ecology or underlying host distributions.

      Page 294: Statements regarding structural similarity appear to rely primarily on visual comparisons. Could the authors provide quantitative structural alignment metrics (e.g., RMSD, TM-score, DALI score, Foldseek score) to support these conclusions?

      Page 321-324: Given the scale and curation of the dataset, the final case-study section is limited. Broader comparative analyses of gene content, ORF architecture, and composition related to host associations, phylogenetic relationships, and/or ecological variables could provide additional evolutionary insights beyond a small number of illustrative examples.

      Page 333: Given the emphasis on the feasibility of petabase-scale sequence mining, the manuscript would benefit from a more detailed description of the computational resources required. The reported ~10-hour runtime is difficult to interpret without information regarding hardware specifications, CPU-hours, memory requirements, storage footprint, and cloud infrastructure (if used). Such information is important for evaluating the reproducibility and practical applicability of the approach.

      Pages 408-409: Were metadata available regarding viral enrichment procedures, particle purification, or size-selection protocols? Such information could influence the interpretation of PV detection frequencies across library types.

      Pages 415-416: The differentiation between viral and endogenous viral sequences is one of the biggest challenges in viral metagenomics and large-scale data mining for viral sequences. This issue is particularly relevant because the distinction between exogenous and endogenous viral sequences may directly affect estimates of novel PV diversity and inferred host associations. The manuscript acknowledges that papillomavirus sequences recovered from DNA-based libraries may derive from integrated viral DNA. However, there is no systematic analysis addressing the potential contribution of endogenous papillomavirus elements (EVEs) to the reported diversity estimates. Given the large number of host genome sequencing projects represented in the SRA, some detected PV-like sequences may correspond to integrated or fossil viral sequences rather than exogenous viruses. The authors could discuss this possibility more explicitly and provide analyses evaluating the prevalence of integration signatures, disrupted ORFs, host-genome flanking regions, or other indicators that would help distinguish endogenous viral elements from actively circulating PVs.

      Pages 532-533: The study is described as "petabase-scale"; however, the analyses were performed on a pre-assembled and compressed representation of the SRA rather than directly on petabase-scale raw sequencing data. The authors may wish to clarify this distinction and explicitly acknowledge that the computational burden is substantially reduced by the Logan framework and, from this perspective of computational power applied, this study is not in the same context as Serratus and Logan.

      Perspective comment: One of the strengths of this study is the generation of a highly curated papillomavirus protein dataset spanning a broad range of known and newly identified PV diversity. Given the increasing importance of structure-based homology detection in virology, the authors may wish to discuss the potential of this resource for future structure-guided discovery efforts. Recent studies have shown that protein structure prediction and comparison can reveal extremely distant evolutionary relationships that are undetectable at the sequence level. The curated PV dataset generated here could serve as a valuable reference for searching unannotated proteins from metagenomic "dark matter" datasets for structural homologs or convergent folds related to papillomavirus proteins. Such approaches may help identify highly divergent PV lineages or previously unrecognized viral proteins that retain structural similarity despite extensive sequence divergence.

    1. eLife Assessment

      This valuable study reports that ALDH-abundant cells exhibit stem cell properties and may play a key role in endometrial epithelial development in mice. The data support the main conclusion and are convincing. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

    2. Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and builds on established hypotheses in the field. The manuscript is well written and scientifically sound and the important experimental limitations and interpretation caveats are presented throughout.

      The majority of my initial comments have been adequately addressed within the text.

      Remaining limitations include:

      (1) The impact of tamoxifen injection directly on Aldh1a1 expression in the developing uterus.

      (2) It would be beneficial to demonstrate the degree of cell death following diphtheria toxin treatment 24-48 hours after injection in Tam-treated mice at PND 10. It is not clear as to why the 4-day timepoint was selected. Cells expressing the DTR should begin undergoing apoptosis within several hours after treatment.

    3. Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      Comments on revised version.

      In the revised manuscript, comments have largely been addressed and the manuscript is improved. The authors have tempered their inference of the contribution of ALDH1A1+ cells to endometrial regeneration, but the conclusions are still somewhat overstated. However, the overall work provides valuable insight and strengthens the growing body of literature characterizing endometrial stem/progenitor cells and their function.

    4. Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      Comments on revised version.

      The authors have answered the questions in the revised manuscript, no further comments.

    5. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This valuable study reports that the ALDH-abundant cells display stem cell properties and may play a key role in the endometrial epithelial development in the mouse. The data supporting the main conclusion are solid, although further improvements are needed to strengthen the conclusions. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

      We thank the reviewers and editor for their critical reading and assessment of our manuscript. We carefully considered each of the points raised by the reviewers. In this document and in the edited manuscript and figures, we have carefully addressed each of the comments and requested modifications. In light of these changes, we expect that you will find that the manuscript has improved.

      We indicate our responses to the reviewers below in blue font and highlight the changes in the manuscript using the line numbers corresponding to the tracked version of the revised document.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and will build on established hypotheses in the field. The manuscript is well written and scientifically sound; however, several experimental limitations and interpretation caveats should be addressed.

      We thank the reviewer for their comments and expert assessment of our paper.

      (1) The methods surrounding the passage number and duration of culture following sorting prior to transcriptomic profiling should be clarified in the figure legends. Related to this, the representative images in Figures 1D and 1E do not appear consistent with the quantification presented in Figures 1F-H and should be reconciled.

      Thanks for this comment. We have now clarified this in the Figure 1 legend as follows,

      LINES 1026-1029: “Organoid formation assay performed immediately after luminal epithelial cell isolation and by plating equal numbers of viable ALDH<sup>LO</sup> (D) and ALDH<sup>HI</sup> (E) epithelial cells. ALDH<sup>LO</sup> and ALDH<sup>HI</sup> organoids were cultured for two weeks and passaged once prior to the organoid formation assays and transcriptomic analyses.”

      Regarding the second comment, we recognize that the images we showed may not have been the most representative of our quantification. As such, we replaced them with the organoid images so that they better reflect the quantification outlined in Figure 1F-H.

      (2) The conclusion that ALDH1A1+ cells are enriched in populations with stem cell characteristics relies primarily on transcriptomic analysis. Protein-level co-localization should be performed to strengthen this claim.

      We thank the reviewer for this comment. Unfortunately, the antibodies for many of these stem cell markers (such as LGR5, AXIN2, and SUSD2) are not well-suited for immunostaining. Others that have been proposed in human and are amenable to immunostaining are not suitable markers for mouse endometrial stem cells (such as CDH2). We hope that by showing that ALDH1A1 is expressed in patterns that are similar to the previously published stem cell markers LGR5 and AXIN2 (i.e., throughout the epithelium in the developing uterus and subsequently enriched in the tips of the endometrial glands of adult mice), along with transcriptomic studies, we can demonstrate its utility as a marker for mouse endometrial stem cells.

      (3) The overlap of 19 genes between the data set here and AXIN2 HI data is presented as evidence of shared stemness identity, but no statistical assessment of this overlap is provided. A hypergeometric test should be performed to determine whether this overlap is greater than expected by chance.

      Thank you for this suggestion. We have performed a hypergeometric test and determined that the reported shared genes between the two datasets are greater than is expected by chance. We have updated the results section to state the following:

      Lines 137-140: "We determined that the overlap between ALDH<sup>HI</sup> and Axin2<sup>+</sup> stemness marker genes was significantly greater than expected by chance for both upregulated (21/346 genes, 1.81-fold enrichment, p = 0.0067) and downregulated (19/674 genes, 1.67-fold enrichment, p = 0.021) gene sets (hypergeometric test, universe = 23,182 genes)."

      (4) The impact of tamoxifen injection on Aldh1a1 expression should be characterized in the neonatal uterus, as tamoxifen itself has known estrogenic activity that could confound interpretation of the lineage tracing results at early postnatal timepoints.

      Although we took measures to control for this possibility by using multiple time-points and models to trace the impact of Aldh1a1<sup>+</sup> cells in development and adulthood, we recognize the importance of this comment and acknowledge that this is a limitation in the design of our study. We have included the following text to the Discussion acknowledging this point:

      Lines 433-441: “Given the well-documented impacts of tamoxifen for lineage tracing studies, it is imperative to use doses of tamoxifen that will minimize estrogenic impacts and result in off-target effects (Rios et al., 2016). This often requires administration at doses that will achieve maximal recombination of the desired gene, while ensuring that the potential deleterious impacts of tamoxifen are minimized (Chen et al., 2023; Pimeisl et al., 2013). The cre/ERT2 tamoxifen inducible model is widely used to study uterine biology where it serves as a useful tool to interrogate the spatiotemporal impact of key genes, either through inactivation or for lineage tracing. Despite its widely documented utility across many tissue types and developmental timepoints, the use of tamoxifen and its impacts on the endometrium remain a limitation of our study, which we tried to address by implementing multiple timepoints, doses, and orthogonal assays in our experimental design.”

      (4b) Related to this, while low-dose tamoxifen is shown to label individual cells within 24 hours of injection, the translation dynamics of the label following Cre-mediated recombination can require up to 72 hours. The presence of only a few labeled clones at PND8 but multiple separate clones per cross-section at later timepoints warrants discussion and may reflect labeling kinetics rather than clonal expansion.

      The reviewer raises an important point. We agree that the 72hr-translation kinetics of the cre-mediated recombination is a legitimate consideration for interpreting our data and we have added the text below to the Discussion section acknowledging this point.

      We have addressed this by adding the following text to the discussion:

      Lines 417-422: We hypothesized that the singly labeled cells observed from one day tracing experiments expanded in a clonal fashion during the various timepoints we measured. We note that the translation kinetics of the labeled cells following cre-mediated recombination may contribute to the limited labeling observed at PND8/PND15 and there is a potential for delayed labeling of cells between 24 and 72 hours of tamoxifen administration. However, the continuous increase in labeled cells at the subsequent timepoints favors our interpretation of clonal expansion as the primary explanation.

      (5) It would strengthen the in vivo ablation data to validate the degree of cell death following diphtheria toxin treatment directly. It is possible that a general decrease in cell number rather than specific loss of a stem cell population is responsible for the observed reduction in gland number and FOXA2 expression (Tongtong et al 2017).

      We agree that this is an important control to incorporate into our experimental design. To rule out this possibility, we performed immunohistochemistry of cleaved caspase 3 in the uterine tissues of DTR<sup>flox/flox</sup> and DTR<sup>flox/flox</sup>;Aldh1a1<sup>cre/ERT2</sup> mice 4 days after administration of diphtheria toxin. The results indicate similar levels of cleaved caspase 3 detection in both genotypes, suggesting that the decrease in FOXA2+ cells is not due to non-specific cell death, but rather the result of ALDH1A1<sup>+</sup> cells. These data and the following text have been added to the manuscript:

      Lines 320-324: “We determined that the decreased in FOXA2<sup>+</sup> cells in the experimental mice was not the result of non-specific DT-mediated cell death, as similar levels of cleaved caspase 3-positive cells were detected in the DT-treated control ROSA26<sup>DTR/DTR</sup> and ROSA26<sup>DTR/DTR</sup>;Aldh1a1<sup>cre/ERT2/+</sup> mice 4 days post-diphtheria toxin administration (Figure S3G-H’).”

      (6) The lineage tracing data in the postpartum endometrium demonstrate that Aldh1a1-marked cells are present during regeneration, but it remains unclear whether these cells are preferentially activated or expanded in response to tissue injury. Coupling these studies with diphtheria toxin-mediated ablation during active regeneration would more directly test the proposed regenerative role of this population.

      This is a great point and one that we would be very interested in pursuing as follow-up studies in our future work. Regretfully, due to the long generation time and experimental procedures associated with these proposed studies, we are not able to include these experiments in the current manuscript. Thus, we have changed our wording and conclusions throughout the manuscript to be less definitive in terms of the role of Aldh1a1 in regeneration, since this will be the focus of future studies.

      The contribution of stromal Aldh1a1 lineage-positive cells is underexplored in the discussion, given the lineage tracing data showing stromal labeling across multiple timepoints and its potential relevance to mesenchymal-to-epithelial transition.

      Thank you for the suggestion. We have now expanded this section in the Discussion to include the following:

      Lines 496-504: We also found ALDH1A1<sup>+</sup> stromal cells were more prevalent when tracing began in adult mice. Other studies have shown that mesenchymal cells contribute to endometrial regeneration in the postpartum phase or after induced menses through a process of MET (Cousins et al., 2014; Kirkwood et al., 2022; Li et al., 2025). Similarly, lineage tracing studies have shown that MET is an active process and contributes to epithelial cell regeneration in the post-partum phase (Huang et al., 2012; Patterson et al., 2013). Although this is an area of active investigation in the field, with some contradicting reports, it is plausible to hypothesize that endometrial tissue has the capacity to undergo wound-healing and regeneration via several mechanisms (Ang et al., 2023; Ghosh et al., 2020). The process of MET in wound healing is widely documented in other organs, such as the kidney, liver and lung, where MET is associated with depletion of the resident epithelial cell pool (Bi et al., 2012; Niayesh-Mehr et al., 2024; Zeisberg et al., 2005).

      Finally, the word 'control' may overstate the functional evidence presented. 'Contribute' may be more accurate given the partial and context-dependent nature of the phenotypes observed.

      We agree with the reviewer’s point that control may overstate the evidence that we provide in the manuscript. To reflect this, we have edited the manuscript title and text to address this suggestion.

      Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      This study addresses an important area of research in the field of endometrial stem/progenitor cell biology. The authors are commended for their use of multiple complementary methods, including lineage tracing, DTR-mediated cell ablation, organoid assays, and RNA-seq in mouse and human models to assess the stem-like nature of Aldh1a1+ cells. The data support the stem/progenitor phenotype of Aldh1a1+ epithelial cells during endometrial development; however, there are noted discrepancies between organoid formation assays and lineage tracing experiments regarding the stemness of Aldh1a1+ epithelial cells in adults. Specifically, organoids were generated from adult cells and demonstrated in vitro stem cell activity; however, in vivo lineage-tracing of adult cells either during the estrous cycle or postpartum repair does not show expansion of Aldh1a1+ cells, suggesting they do not have stem/progenitor activity. Additionally, the stem-ness of epithelial vs stromal Aldh1a1+ cells is confounded in the study because epithelial cells were not purified for organoid experiments, epithelial cells were not exclusively lineage-traced as stromal cells were also labeled, and mesenchymal-epithelial transition was suggested to occur during postpartum repair. The following specific comments are presented to detail these concerns:

      We thank the reviewer for their critical reading of our manuscript and constructive comments.

      (1) The statement in the brief summary, "...critical for lifelong endometrial regeneration," is not supported by the data provided.

      We have edited the brief summary to exclude this statement, it now reads as follows:

      Lines 4-5: “We uncover ALDH1A1<sup>+</sup> cells as a group of hormone sensitive stem cells contributing to endometrial development and regeneration.”

      (2) AlDH1A1 is not restricted to the endometrial epithelium, and epithelial cells were not purified by flow cytometry for experiments in Figure 1. Figure 2 clearly shows the presence of mesenchymal cells, even using the described method for enriching for epithelial cells. Therefore, contaminating mesenchymal cells with high ALDH activity may confound the experimental results in Figure 1, either through promoting epithelial cell growth or through MET. The authors should provide clear evidence of epithelial purity in organoid experiments or that mesenchymal cells are not contained in the ALDHhi population. These comments also apply to the human organoid experiments in Figure 7.

      We thank the reviewer for raising this important point. Our group has been using the enzymatic method to routinely separate epithelial from stromal cell populations from the mouse uterus (see references dating back to 2015, PMID 26721398, 28324064, 34099644). In these experiments we typically obtain >98% purity in the epithelial and stromal cell compartments, respectively. We can directly observe this purity in the immunofluorescence images shown I Author response image 1 and Author response image 2, where mouse endometrial epithelial cells and stromal cells were enzymatically separated and immunostained with E-cadherin and vimentin antibodies to detect epithelial and mesenchymal cells in both cell preparations. The images show very few contaminating epithelial and stromal cells in either cell preparation. We have observed similar results when preparing epithelial and stromal cell preparation from the human endometrium, where the epithelial cell organoids display high purity with ~100% epithelial cell expression when we perform immunostaining.

      Author response image 1.

      Purity of mouse endometrial epithelial cells obtained via enzymatic and mechanical dissociation. A-B) Shows the epithelial (A) and stromal (B) cells plated on glass coverslips and immunostained with an epithelial cell marker (cytokeratin 8, red), a stromal cell marker (vimentin, green), and DAPI.

      Author response image 2.

      Human endometrial epithelial organoids were fixed and immunostained with cytokeratin 8 (green) and DAPI. The images are typical for our epithelial cell cultures and demonstrate that all epithelial cells are CK8-positive.

      (3) Lines 186-187: Susd2 was increased in EpSC clusters, yet this is a mesenchymal stem/progenitor marker in humans. The authors should discuss the implications of this.

      We thank the reviewer for highlighting this. We have now included the following in our Discussion to address this point:

      Lines 527-532: Clustering with this population of EpSCs were Susd2<sup>+</sup> cells, which are well-characterized mesenchymal progenitors that are enriched in the perivascular regions of the human endometrium (Darzi et al., 2016; Khanmohammadi et al., 2021). The presence of Susd2<sup>+</sup> cells, while unexpected in an epithelial stem cell niche, could indicate the presence of a transitional mesenchymal or perivascular cell that is differentiating into epithelium. Evidence for both mesenchymal and Nestin2<sup>+</sup> pericytes have been recently described in the mouse endometrial epithelium (Kirkwood et al., 2022; Li et al., 2025).

      (4) In Figure 5, RFP+ epithelial cells should be quantified as in previous figures to substantiate the statement in lines 279-280, "At PPD5, the proportion of RFP+ epithelial cells had expanded relative to PPD1 and PPD3 (Figure 5E-E')." Especially because in the low mag images (C-E), RFP+ epithelial cells appear to be most abundant at PPD1 and decrease at PPD3 and PPD5, suggesting that they may not be involved in endometrial regeneration/repair (contradicting the interpretation in line 285). Further, if there is in fact a decrease over postpartum repair, then regeneration should be removed from the title of the manuscript. RFP+ stromal cells should also be quantified.

      We appreciate this reviewer’s comment and agree that as stated, the conclusion is not fully supported by the data. To address this comment, we have edited the results so that they clearly indicate the results and remove any ambiguity:

      As requested, we quantified the number of RFP+ stromal and epithelial cells during the postpartum phase and noted that RFP+ cells were prominent in the stromal compartment of the endometrium. While RFP+ epithelial were also observed during these timepoints, they were less abundant than RFP+ stromal cells. Because the number of RFP+ cells did not significantly change over the postpartum phases in neither the stromal nor epithelial compartment, we have modified our conclusion to state that ALDH1A1+ cells are transiently detected in the regenerating endometrium.

      Results:

      Lines 287-294: “By analyzing the uterine tissues near the placental detachment site, we observed that RFP positive cells were prominent in the endometrial stromal cells that were adjacent to the luminal epithelium (Figure 5C-C’, green arrows). RFP<sup>+</sup> cells were also observed in the stromal cells near the placental detachment sites at PPD1 and PPD3 (Figure 5D’-E’, red & blue arrows) and in limited luminal epithelial cells (Figure 5D”,E”). Quantification of RFP+ cells throughout these postpartum phases indicated that stromal cells had more frequent ALDH1A1<sup>+</sup> stromal cells (360 ± 103, PPD1, n=3; 217 ± 107, PPD3, n=3; 254 ± 32, PPD5, n=4) than ALDH1A1<sup>+</sup> epithelial cells in the regenerating endometrium (65 ± 65, PPD1, n=3; 20 ± 10, PPD3, n=3; 114.25 ± 39, PPD5, n=4) (Figure S4).”

      Discussion:

      Lines 512-520: “We also noted that a majority of ALDH1A1<sup>+</sup> cells were localized to the active areas of endometrial regeneration near the placental detachment sites at PPD1 with a pronounced expression in the sub-epithelial stromal cells. As regeneration progressed, we continued to observe ALDH1A1<sup>+</sup> cells in the stromal compartment within the placental detachment sites at PPD3 and PPD5, with a progressive, but not statistically significant, increase in ALDH1A1<sup>+</sup> epithelial cells. Collectively, our data demonstrate that ALDH1A1<sup>+</sup> lineage cells participate in the restoration of endometrial architecture and functional compartments in the postpartum phase, even if their direct contribution is transient. Future detailed and mechanistic studies will be necessary to fully characterize their role in this process and their long-term consequence in postpartum regeneration.”

      (5) For Figure 7F, it should be clearly stated in the main text that the results are from one patient sample and the data presented are experimental replicates, so as not to be confused with biological replicates (the same for Supplementary Figure S4). Were B and G in Figure 7 also from one patient?

      Thanks for pointing this out. We have edited the figure legends in the main text and supplemental figures to indicate this.

      Lines 336-337: “…main figures show representative results from one patient sample performed in technical replicates, with additional patient samples included in the supplement…”

      (6) Lines 425-427: "Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed a complete restriction of ALDH1A1 to the glandular crypts." In Figure 2 S' ALDH1A1+ cells are visible in the LE (the staining is lighter than in the GE but looks real), contradicting this statement.

      This is an important distinction. We have now edited this part of the manuscript to state:

      Lines 458-461: “Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed enriched ALDH1A1 in the glandular crypts with weak luminal epithelial staining, while the ovariectomized controls had strong ALDH1A1 expression throughout the luminal and glandular epithelium.”

      (7) Lines 466-467: "In cycling mice, we found sporadic cells that expressed both stromal and epithelial markers in the ALDHA1+ cells." These data are not presented.

      We apologize for the confusion, this sentence has been removed from the discussion.

      (8) These data support the role of Aldh1a1+ cells in endometrial epithelial development, but conclusions about their role in repair/regeneration should be tempered as the data are much weaker here.

      We thank the reviewer for their overall assessment. To address this point, we have thoroughly edited the appropriate areas to temper the conclusions and ensure that they are strongly supported by our data. We have also edited the manuscript’s title to reflect this.

      Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      We thank the reviewer for their assessment of our manuscript.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      We thank the reviewer for their expert assessment and positive comments regarding our manuscript.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As this study evaluated Aldh1a1+ cells in both the epithelium and stroma, it is recommended that the title and abstract be revised to reflect this.

      Thank you. Both the title and abstract title have been updated to reflect this comment.

      (2) Lines 46-47 in the abstract: "Aldh1a1+ cells expanded during postnatal development, estrus cycling, and following post-partum repair." It is recommended to clarify stromal vs epithelial Aldh1a1+ cell expansion, as only the stromal cells expanded during the estrous cycle. Also, RFP+ epithelial cells were not quantified during postpartum repair and visually appear to decrease (see comment below regarding Figure 5), so this statement is misleading.

      The abstract was edited following the suggestions so that it depicts our results. Similarly, we have addressed the comment regarding Figure 5 and the interpretation of the post-partum regeneration experiments (see comment above for the full explanation and edits to the manuscript).

      (3) Lines 65-69, 186-187: the authors should be clear when describing markers of putative epithelial vs mesenchymal stem/progenitor cells in the introduction (e.g., CDH2/SSEA1/SOX9 for epithelial and SUSD2 for mesenchymal).

      Thank you for this suggestion. The Introduction section now states the following:

      Lines 70-73: “These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).”

      (4) Lines 78-80: "Studies tracing the fate, ablation, and proliferative capacity of Lgr5+ cells in the uterus identified an Lgr5+ niche that is enriched in the crypts of the glandular epithelium and promotes endometrial regeneration (Seishima et al., 2019)." This statement is incorrect regarding regeneration, as Lgr5 marks stem/progenitor cells in the developing uterus but not the adult, during which endometrial regeneration occurs. Please revise.

      Thank for you for this clarification. The statement has been revised and now reads as follows:

      Lines 82-84: “Studies tracing the fate, ablation, and proliferative capacity of Lgr5<sup>+</sup> cells in the uterus identified an Lgr5<sup>+</sup> niche that marked stem/progenitor cells in the developing uterus (Seishima et al., 2019).”

      (5) Recommend using free-form shapes to outline GE, LE, and EpSC in Fig 2C and stating in the text which clusters correspond to LE and GE (lines 167-169).

      Thank you, this has been edited in Figure 2C and in the text, which now reads as follows:

      Lines 176-177: “Clusters 5, 7, 21 and 16 were classified as glandular epithelial cells, and clusters 0, 2, 3, 10, 14, and 24 were classified as luminal epithelial cells (Figure 2C).”

      (6) "Estrus" refers to the specific stage of the "estrous" cycle. Estrus cycle is incorrect and should be estrous cycle.

      Thank you for pointing this out. It has been corrected throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Suggest increasing the font size for some of the labels in the figures.

      Thank you, we have increased font size in the figures to improve the quality.

      (2) Need to include more details of the human endometrial tissue used in this study: pathology, age, and stage of menstrual cycle.

      We thank the reviewer for the helpful suggestion. The details that are available to us have been included in Supplementary Table S5.

      (3) Line 65-67 - Different endometrial stem cell subsets - SUSD2+ reside in perivascular regions, while other markers are located in glandular epithelium, need revision.

      Thank you, this has been revised in the Introduction. The area now reads as follows:

      Lines 70-73: These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).

      (4) Include a discussion about the interpretation of their current finding in relation to the dynamic regeneration observed in human endometrium due to menstrual bleeding/tissue breakdown, compared to the cycles of growth and regression that occur in mice.

      This is a great suggestion. We have added the following statement to our Discussion section:

      Lines 537-544: Additionally, our studies in human endometrium extend our characterization of ALDH1A1 as an adult endometrial stem cell marker and emphasize the importance of ALDH1A1+ in the regenerative potential of the endometrium. The conserved hormonal responses between human and mouse endometrium support the hypothesis that cycles of proliferation, differentiation, and regression, whether through resorption/autophagy in mice or menstrual breakdown in humans, are governed by concerted growth factor signaling and dedicated stem cell populations with the capacity to expand and differentiate across repeated cycles of repair. Our detailed studies in both mouse and human models indicate that ALDH1A1+ cells represent a dedicated cell type within the endometrium with the potential to drive repair during adulthood. Collectively, these findings advance our understanding of the mechanisms that control endometrial cycling and regeneration throughout the reproductive lifespan.

      Reference

      Ang, C.J., Skokan, T.D., and McKinley, K.L. (2023). Mechanisms of Regeneration and Fibrosis in the Endometrium. Annu Rev Cell Dev Biol 39, 197-221.

      Bi, W.R., Jin, C.X., Xu, G.T., and Yang, C.Q. (2012). Bone morphogenetic protein-7 regulates Snail signaling in carbon tetrachloride-induced fibrosis in the rat liver. Exp Ther Med 4, 1022-1026.

      Chen, M.Y., Zhao, F.L., Chu, W.L., Bai, M.R., and Zhang, D.M. (2023). A review of tamoxifen administration regimen optimization for Cre/loxp system in mouse bone study. Biomed Pharmacother 165, 115045.

      Cousins, F.L., Murray, A., Esnal, A., Gibson, D.A., Critchley, H.O., and Saunders, P.T. (2014). Evidence from a mouse model that epithelial cell migration and mesenchymal-epithelial transition contribute to rapid restoration of uterine tissue integrity during menstruation. PLoS One 9, e86378.

      Cousins, F.L., Pandoy, R., Jin, S., and Gargett, C.E. (2021). The Elusive Endometrial Epithelial Stem/Progenitor Cells. Front Cell Dev Biol 9, 640319.

      Darzi, S., Werkmeister, J.A., Deane, J.A., and Gargett, C.E. (2016). Identification and Characterization of Human Endometrial Mesenchymal Stem/Stromal Cells and Their Potential for Cellular Therapy. Stem Cells Transl Med 5, 1127-1132.

      Ghosh, A., Syed, S.M., Kumar, M., Carpenter, T.J., Teixeira, J.M., Houairia, N., Negi, S., and Tanwar, P.S. (2020). In Vivo Cell Fate Tracing Provides No Evidence for Mesenchymal to Epithelial Transition in Adult Fallopian Tube and Uterus. Cell Rep 31, 107631.

      Huang, C.C., Orvis, G.D., Wang, Y., and Behringer, R.R. (2012). Stromal-to-epithelial transition during postpartum endometrial regeneration. PLoS One 7, e44285.

      Khanmohammadi, M., Mukherjee, S., Darzi, S., Paul, K., Werkmeister, J.A., Cousins, F.L., and Gargett, C.E. (2021). Identification and characterisation of maternal perivascular SUSD2(+) placental mesenchymal stem/stromal cells. Cell Tissue Res 385, 803-815.

      Kirkwood, P.M., Gibson, D.A., Shaw, I., Dobie, R., Kelepouri, O., Henderson, N.C., and Saunders, P.T.K. (2022). Single-cell RNA sequencing and lineage tracing confirm mesenchyme to epithelial transformation (MET) contributes to repair of the endometrium at menstruation. Elife 11.

      Li, S.Y., Whiteside, S., Li, B., Sun, X., and DeFalco, T. (2025). Mesenchymal-to-epithelial transition of perivascular cells contributes to endometrial re-epithelialization. Nat Commun 16, 10174.

      Niayesh-Mehr, R., Kalantar, M., Bontempi, G., Montaldo, C., Ebrahimi, S., Allameh, A., Babaei, G., Seif, F., and Strippoli, R. (2024). The role of epithelial-mesenchymal transition in pulmonary fibrosis: lessons from idiopathic pulmonary fibrosis and COVID-19. Cell Commun Signal 22, 542.

      Patterson, A.L., Zhang, L., Arango, N.A., Teixeira, J., and Pru, J.K. (2013). Mesenchymal-to-epithelial transition contributes to endometrial regeneration following natural and artificial decidualization. Stem Cells Dev 22, 964-974.

      Pimeisl, I.M., Tanriver, Y., Daza, R.A., Vauti, F., Hevner, R.F., Arnold, H.H., and Arnold, S.J. (2013). Generation and characterization of a tamoxifen-inducible Eomes(CreER) mouse line. Genesis 51, 725-733.

      Rios, A.C., Fu, N.Y., Cursons, J., Lindeman, G.J., and Visvader, J.E. (2016). The complexities and caveats of lineage tracing in the mammary gland. Breast Cancer Res 18, 116.

      Seishima, R., Leung, C., Yada, S., Murad, K.B.A., Tan, L.T., Hajamohideen, A., Tan, S.H., Itoh, H., Murakami, K., Ishida, Y., et al. (2019). Neonatal Wnt-dependent Lgr5 positive stem cells are essential for uterine gland development. Nat Commun 10, 5378.

      Zeisberg, M., Shah, A.A., and Kalluri, R. (2005). Bone morphogenic protein-7 induces mesenchymal to epithelial transition in adult renal fibroblasts and facilitates regeneration of injured kidney. J Biol Chem 280, 8094-8100.

    1. eLife Assessment

      The authors use solid high resolution microscopy techniques to present a model in which SARS-CoV-2 attachment and endocytosis are mediated by heparan sulfate, whereas ACE2 only functions downstream of these early processes. These findings serve as a potentially valuable starting point to examine SARS-CoV-2 entry models that challenge current paradigms. However, examination of the model in a clean heparan sulfate deficient background and in the context of TMPRSS2-dependent plasma membrane fusion are lacking. Thus, the broader impact remains uncertain in the absence of additional controls and orthogonal approaches.

    2. Reviewer #1 (Public review):

      This revised paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. The authors used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions downstream of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. The authors conclude their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and use of primary airway cells. While some studies were performed in the revision to address the Reviewer concerns, which improved the paper clarity, others were cursorily addressed, which limit the impact of the studies. Particularly. experiments were not performed to account for TMPRSS2 expression and plasma membrane fusion. Moreover, addition of studies in which hACE2 is expressed in cells genetically lacking HS were not designed. Thus, it the picture remains unclear picture exactly where downstream hACE2 functions and how this might differ given new structural models of TMPRSS2 activation (PMID: 42050172), which occur after ACE2 recognition of spike on the cell surface.

    3. Reviewer #2 (Public review):

      In the manuscript by Han et al, the authors assess binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that SARS-CoV-2 spike (on the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well described but the authors here attempt to make the claim that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. Additional controls would be of great benefit.

      Further, it is this reviewers opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. A more balanced and nuanced view of their interesting data would be of value.

      Major:

      The authors need to rigorously define a "receptor." This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm) while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is not necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this definition. This is proven in Fig 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959 but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. This is inconsistent with the claim that HS is broadly important (beyond the BHK cells overexpressing ACE2 that are used here).

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide inadequate discussion into the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2 deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used as protease expression is a key determinant of the site of fusion.

      An alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays are TMPRSS2 vs Cathepsin L dependent (using E64d and camostat for instance) as mentioned above.

      Each figure legend should clearly state how many independent experiments and replicates per experiment were performed.

      All bar plots should show individual dots (i.e. Fig 1G) to better reveal the variance of each dataset.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. A prevailing model that has been developed over the last five years by work from many laboratories using a variety of biochemical, structural, and microscopic approaches is that HS acts a co-receptor for SARS-CoV-2; its binding to SARS-CoV-2 both concentrates virus on the surface of target cells and allosterically alters the spike protein to promote an "up/open" RBD conformation that enables engagement of the proteinaceous receptor human ACE2 on the cell surface (PMID: 32970989, 35926454, 38055954, 39401361, 40548749). These two events enable plasma membrane fusion (after a cleavage event promoted by plasma membrane TMPSS2) or endocytosis and subsequent pH-dependent fusion (which requires a cathepsin L-mediated cleavage of the spike).

      The authors in this study used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions "downstream" of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. Blocking HS binding with pixantrone, a drug under clinical evaluation for cancer (due to its anti-topoisomerase II activity), inhibited SARS-CoV-2 Omicron JN.1 variant from attaching to and infecting human airway cells. The authors conclude that their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the interesting and technically well-executed study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and the use of primary airway cells. Nonetheless, there are issues that need to be addressed to buttress the proposed model compared to earlier ones. These include: (a) the distinction between macropinocytosis and receptor-mediated endocytosis and what this might mean for productive SARS-CoV-2 infection; (b) the need to account for TMPRSS2 expression and plasma membrane fusion; (c) addition of genetic studies in which hACE2 is expressed in cells lacking HS; (d) an unclear picture of exactly where downstream hACE2 functions; and (e) and a need for comparative/additional study of earlier SARS-CoV-2 variants, which preferentially fuse at the plasma membrane.

      We thank the reviewer for the strong support of this manuscript. We addressed the reviewer’s concerns in the Recommendations to the authors. We did not distinguish whether the endocytic route is macropinocytosis or receptor-mediated endocytosis, because it is a separate study beyond the scope of the present work. We did not examine earlier SARS-CoV-2 variants because we considered it a study beyond the scope of the present work, but a good idea that we may work on in the future. For detail on how we address the remaining concerns, please see our response to the reviewer’s Recommendations for the authors.

      Reviewer #2 (Public review):

      In this manuscript by Han et al, the authors assess the binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that the SARS-CoV-2 spike (in the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface, which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well-described, but the authors attempt to make the claim here that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears to be of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. The manuscript is imprecise and far overstates the actual findings shown by the data. Additional controls would be of great benefit.

      Further, it is this reviewer's opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. While the manuscript provides new insight into the details of this process, the manuscript attempts to oversell this finding by applying new words rather than new molecular details. The authors would be better served by presenting a more balanced and nuanced view of their interesting data. In this reviewer's opinion, the salesmanship significantly detracts from the data and manuscript.

      We thank the reviewer for pointing out that our manuscript is of interest to the field. However, we do not think that we oversell our data. hACE2 has been widely considered the receptor (or the binding partner) that mediates SARS-CoV-2 cell-surface attachment, whereas HS is considered only an attachment factor that facilitates SARS-CoV-2 binding with hACE2 at the cell surface. In the present work, we found that HS, but not hACE2, is the cell-surface attachment receptor (or binding partner), whereas hACE2 is not essential for attachment, but acts downstream of virus endocytosis to facilitate viral genome expression. This finding suggests significant modification of the current model by replacing the attachment receptor (or binding partner) from hACE2 to HS, treating HS as a primary receptor rather than an attachment factor, and relocating the hACE2 action site from the cell surface to the endosome. For these reasons, we do not consider these statements overselling our data. However, as the reviewer suggested in his/her specific comments, we revised the manuscript to ensure that we did not overgeneralize our findings (see our responses to the reviewer’s Recommendations to the authors).

      Major Comments:

      The authors need to rigorously define a "receptor" vs an "attachment factor." They also should avoid ambiguous terms such as "receptor underlying ...attachment" and "attachment receptor" (or at least clearly define them). Much of their argument hinges on the specific definition of these terms. This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm), while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is neither necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this fairly textbook definition. This is proven in Figure 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959, but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. The authors need to discuss this important information and reconcile it with their data and model if they want to claim that HS is broadly important.

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction, which likely won't excite a drug company. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide little discussion of the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2-deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used, as protease expression is a key determinant of the site of fusion.

      The claim that "SARS-CoV2 JN.1 variant binds to heparan sulfate, not hACE2, in primary human airway cells" is extraordinary and thus requires extraordinary evidence.

      First, PIX reduces attachment by 5-fold, which is not the same as "nearly abolished." Also, anti-ACE2 "nearly abolished" entry in 7D, while PIX did not. If the authors want to make these claims, an alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays is TMPRSS2 vs Cathepsin L dependent (using E64d and camostat, for instance) as mentioned above.

      Each figure should clearly state how many independent experiments and replicates per experiment were performed. What does "3 experiments" mean? Are these three independent experiments or three wells on one day?

      In the well-accepted current model, hACE2 is considered the receptor mediating SARS-CoV-2 cell-surface attachment, entry into cells, and infection, whereas HS is an attachment factor that facilitates SARS-CoV-2 binding to hACE2 at the cell surface. The present work revises this view: HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, with ACE2 acting downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined.

      We made this point clearer throughout the newly revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docks at HS clusters.

      The cited CRISPR-screen literature supports context-dependent host-factor usage. However, the absence of HS biosynthesis genes from a given screen does not prove that HS is irrelevant in that cell type; it only indicates that HS biosynthesis was not detected as a genetic dependency under that assay’s conditions. Such negative results can reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency. In the revised manuscript, we included the following in the Discussion:

      “While some studies using genome-wide CRISPR screening to identify genes involved in SARS-CoV-2 reveal genes for HS biosynthesis, others do not (45-50). The negative result, which might reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency, needs to be verified with specific gene knockout.”

      The ~5-fold reduction is likely due to the inhibitor not completely abolishing HS-virus binding. We revised the Discussion to strengthen the suggestion that targeting the virus cell-surface attachment by interfering HS binding is a therapeutic strategy to prevent and treat COVID-19, as in the following:

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding. However, this strategy has not been the focus for developing methods to prevent and treat COVID-19, likely because HS is considered only a regulator that is not essential for SARS-COV-2 entry. Our finding that HS is the attachment receptor re-emphasizes the importance of perturbing virus-HS binding, the first step of the viral entry, to efficiently block SARS-CoV-2 infection. Further supporting this view, inhibition of HS binding with a clinically used HS-binding agent, pixantrone, inhibits authentic SARS-CoV-2 JN.1 subvariant binding with HS on the cell surface and infection in primary human airway cells (Figs. 6, 7). These results suggest a combinatorial anti-SARS-CoV-2 strategy: early HS blockade to prevent attachment combined with ACE2 targeting to inhibit post-attachment steps”

      We include a sentence in the Discussion that our suggestions are limited to the cells we examined as below.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      Three experiments refer to three independent experiments. We added “independent” accordingly throughout the manuscript.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors define a new paradigm for the attachment and endocytosis of SARS-CoV-2 in which cell surface heparan sulfate (HS) is the primary receptor, with ACE2 having a downstream role within endocytic vesicles. This has implications for the importance of targeting virion-HS interactions as a therapeutic strategy.

      Strengths:

      The authors show that viruses are internalized via dynamin-dependent endocytosis and that endocytic internalization is the major pathway for pseudotyped SARS-CoV-2 genome expression. They show that HS-mediated viral attachment is a critical step preceding viral endocytosis and also subsequent genome expression. Further, they show that hACE2 acts downstream of endocytosis to promote viral infection, and may be co-internalised with virions after HS attachment. Pseudotyped virus and authentic SARS-CoV-2 provide similar results. In addition, the authors demonstrate that remarkable clusters of multiple HS chains exist on the cell surface, visualised by a number of elegant microscopy methods, and that these represent the docking sites for virions. These visualisations are an important general contribution in themselves to understanding the nanoscale interactions of HS at the cell surface.

      The use of a complementary range of methods, virus constructs, and cell models is a strength, and the results clearly support the conclusions.

      Overall, the results convincingly demonstrate a different model to the currently accepted mechanism in which the ACE2 protein is regarded as the cell surface receptor for SARS-CoV-2. Here, the authors provide compelling evidence that cell surface clusters of HS are the primary docking site, with ACE2 interactions occurring later, after endocytosis (whilst still being essential for viral genome expression). This is an exciting and important landmark evidence which supports the view that HS-virion interactions should be viewed as a key site for anti-viral drug targeting, likely in strategies that also target the downstream ACE2-based mechanism of viral entry within endosomes.

      We thank the reviewer for the strong support of the present work.

      Weaknesses:

      This reviewer identified only minor points regarding citing and discussing other studies and typos, which can be corrected.

      We have addressed these points in the revised manuscript. For detail, please see our response to the reviewer’s Recommendations to the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Pathway of internalization.

      The authors show clearly that labeled SARS-CoV-2 (pseudovirus or authentic virus) can become internalized in cells lacking hACE2, and this process depends on HS. However, they also show that this pathway is non-productive with regard to infection. Are the entry vesicles mediated by HS alone, HS + hACE2, and hACE2 alone the same? Or does the combination of co-receptor (HS + hACE2) drive SARS-CoV-2 into endocytic vesicles, whereas HS alone promotes macro- or micropinocytosis (lines 361-362). If HS alone directed SARS-CoV-2 into a non-productive entry vesicle, then hACE2 likely would be acting concurrently with HS and not downstream. A more detailed analysis of the different entry vesicles/pathways that occur with HS alone, HS + hACE2, and hACE2 alone is needed.

      During endocytosis, we did not detect a difference in the size distribution of virus-containing vesicles between BHK (HS alone) and BHK<sub>hACE2</sub> cells (HS+hACE2) (Fig. 2D). The similarity in the vesicle size suggests a similar endocytic path with HS alone or with HS + hACE2. In the revised manuscript, we added the following sentence.

      “Third, 3D-STED imaging showed that A490-labeled vesicle’s full-width-at-half-maximum (W<sub>H</sub>) was 363 ± 17 nm (n = 55) in BHK cells, similar to that (333 ± 13 nm, n = 70) in BHK<sub>hACE2</sub> cells (Fig. 2B-D), supporting a similar endocytic path regardless of hACE2 presence or not.”

      (2) TMPRSS2 and plasma membrane fusion.

      Although the authors allude to membrane fusion as an alternate mechanism of entry, their mechanistic experiments do not address the roles of HS and hACE2 in this process, possibly because their BHK and other cells do not co-express significant levels of TMPRSS2. While many Omicron variants preferentially enter cells via endocytosis (relative to antecedent strains in the pandemic) because of spike mutations that reduce cleavage by TMPRSS2 (PMID: 35104837, 36625591, 35145066), plasma membrane fusion can still occur. The authors should add experiments with co-expression of TMPRSS2/hACE2 [with or without HS] and earlier SARS-CoV-2 variants to establish the role of HS in plasma membrane fusion. Also, are there differences in entry pathways if viruses are prepared in cells expressing TMPRSS2?

      We thank the reviewer for this important comment and agree that our mechanistic experiments were not designed to address TMPRSS2-supported plasma membrane fusion. The reviewer’s suggestion for direct testing of HS function in TMPRSS2-supported plasma membrane fusion, including hACE2/TMPRSS2 co-expression and comparison with earlier SARS-CoV-2 variants, will require dedicated experiments and detection of the fusion pathway that we have not yet designed. It is beyond the scope of the present work. In the revised manuscript, we clarify that the observed ACE2-independent uptake and the predominance of endocytic entry refer to the tested cell systems and do not exclude TMPRSS2-dependent plasma membrane fusion in other cell types, as in the following.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Experiments with hACE2 in cells lacking HS.

      Apart from drug treatment (heparinases or pixantrone) studies shown, the current studies do not directly address whether expression of hACE2 on human cells can allow for endocytosis and productive infection in the complete and genetic absence of HS. The only experiments that use genetically deficient cells are the CHO [hamster] cell studies, and these cells lack hACE2 expression. The authors should knock out a key HS biosynthesis gene (e.g., B4GALT7) in more relevant human cells (e.g., A549-hACE2; ideally with sorted subpopulations having different levels of surface hACE2 expression) and assess endocytosis and infection. This is important given studies in the literature by others suggesting that KO of HS expression reduces but does not abrogate SARS-CoV-2 infection.

      We thank the reviewer for these comments. We showed that virus endocytosis is independent of hACE2 (Fig. 1). The reviewer’s question is whether hACE2 alone can allow for endocytosis of viruses. We have shown that in either BHK (without hACE2) or BHK<sub>hACE2</sub> cells (BHK cells expressed with hACE2), heparinase I/II/III mixture (HPRase) nearly abolished cell-surface immunolabelled HS (Fig. 3D), reduced cell-surface virus attachment by ~83-85% (Fig. 3E), reduced viral uptake by ~80% (Fig. 3F, 3G). These results suggest that hACE2 is not essential for viral attachment and endocytosis. We did not test whether hACE2 alone (without HS) plays a minor role for viral attachment and endocytosis, because to our knowledge, HS is present in nearly every cell. Under this physiological condition, it is HS, not hACE2, that plays an essential role in viral cell-surface attachment and endocytosis. In the revised manuscript, we added a sentence admitting that we did not test whether hACE2 alone is sufficient to support viral uptake and productive infection, as in the following.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis.

      (4) hACE2 function in entry.

      In many places, the authors suggest that hACE2-spike functional interaction occurs "downstream" of HS-dependent binding and endocytosis (e.g., lines 25, 33, 48, 210, 309, 312, 318, 333, 346). However, in their model, it is not clear where exactly this interaction occurs. Are the authors suggesting that this spike binds hACE2 on the cell surface, but this has nothing to do with endocytosis, or that the interaction with hACE2 is occurring at a post-entry step? Can they experimentally demonstrate the stage at which hACE2 is functioning? Is it the same or different in cells lacking HS? What about when TMPRSS2 is present?

      We showed that viral attachment and endocytosis are independent of hACE2, whereas entry as determined by viral gene expression, depends on hACE2. We also showed that most virions bind to HS, not hACE on the cell surface. Based on these results, we propose a model that hACE2 functions downstream of virion endocytosis. We cannot rule out the possibility that a small subset of viruses can also bind to hACE2 after their binding with HS at the cell surface.

      We have not been able to design an experiment to visualize hACE2 mediated virion fusion in endosomes, where hACE2 may facilitate virus fusion and delivery of viral genomes to the cytosol. Productive infection requires only a limited number of successful virion–hACE2 engagement events. While many internalized virions can be visualized, the specific virion or vesicle that ultimately gives rise to productive infection cannot be identified from the present imaging data. This makes it difficult to trace the precise stage or compartment in which the functionally relevant spike–hACE2 interaction occurs. In the revised manuscript, we added a paragraph discussing this limitation as below.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis. Although our data suggest that hACE2 functions downstream of endocytosis to facilitate viral fusion at the endosome for genome delivery to the cytosol, we do not know whether hACE2 binding with the virus occurs at the cell surface or endosomes. The binding may occur in both places, but not essential for virus attachment and endocytosis.”

      (5) Other comments.

      (a) Figure 1A. "Antibody" is misspelled.

      Corrected. Thank you.

      (b) The imaging experiments with pseudoviruses and authentic viruses lack any information on the multiplicity of infection or the number of virions added per cell. If this is particularly high and non-physiological (e.g., >100), is it possible that such conditions might enable viruses to enter [dominantly] through secondary [non-infectious] pathways?

      To address the reviewer’s concern, we used flow cytometry to measure cell-associated VSV-S signal as we diluted the virus by ~600-fold. We found that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHK<sub>hACE2</sub> cells over a ~600-fold dilution of the virus (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations. In the revised manuscript, we included the following sentence and Fig. S7 (Supplementary Information).

      “Flow cytometry also showed that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHKhACE2 cells over a ~600-fold dilution of the virus concentration (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations.”

      (c) Figure 1C and elsewhere. Most of the internalization studies rely on various imaging modalities to demonstrate the pseudovirus or virus on or in the cell. The experiments would be strengthened by inclusion of data from orthogonal binding/internalization assays that measuring virion-associated viral RNA on the surface [4oC binding assay] or inside the cell [after a 37oC temperature shift and exogenous proteinase K and RNAse A treatment]) - such assays can be performed at much lower MOI (e.g., <1, addressed comment #2 above) an also allow more objective quantitation and kinetic analyses of virus internalization (e.g., 0, 5, 15, 30 min at 37oC).

      We demonstrate virion attachment and uptake using multiple approaches, including confocal, STED, and EM analysis, showing virions with the expected morphology at the cell surface and in the cytosol. Furthermore, flow cytometric analysis provides population-level quantitation supporting the same overall conclusion. Thus, while we appreciate and agree that an RNA-based binding/internalization assay would provide additional information, we do not consider it essential to the main conclusion of this work.

      (d) Figure 2. (i) Is there any indication of which vesicles the bath dye is in? Is most of this fluid taken up by micropinocytosis? Are these the same vesicles where the virus that is destined for productive infection (HS/hACE2 engaging) transits? (ii) In all panels, can the authors clearly indicate/label which cells are being used (BHK or BHK-hACE2)? (iii) For the studies with dynasore or dominant-negative dynamin-2-K44A, the readout is at 24 h, a late timepoint, which also could affect virus egress and spread. Can the studies be repeated at much earlier time points (e.g., 15 min to 2 h) to demonstrate that viruses are internalized via dynamin-dependent endocytosis in these cells?

      (i) The bath dye A490 was used as a fluid-phase marker for endocytic uptake, rather than as a marker for a specific vesicle class or intracellular compartment. In principle, any vesicle that takes up extracellular fluid could become labelled by this approach. Since nearly all viruses are in the A490-containing vesicles, productive virus infection must come from some of these vesicles.

      (ii) In the revised Fig. 2 legends, we explicitly indicate which cells are used for each panel.

      (iii) To address the reviewer’s concern, we examined earlier time points for dynasore treatment and found that the virus uptake and genome expression were already markedly reduced at 1 h and 8 h after virus incubation. In the revised manuscript, we described these results as below and in Fig. S5.

      “Fourth, dynasore or dominant-negative dynamin 2-K44A overexpression, which inhibits fission of dynamin-dependent endocytosis [28-30], substantially reduced V-A647 internalized 1-24 h after viral incubation (Figs. 2F-G, S5).

      In addition to inhibiting V-A647 endocytosis, dynasore or dynamin 2-K44A inhibited V-EGFP expression 8-24 h after virus incubation by ~66-77% (Figs. 2F-G, S5), suggesting that endocytosis is the main route for viral genome expression.”

      (e) Line 225. "Envelop" should be "envelope".

      Corrected, thank you.

      (f) Line 235. The authors should clarify that they conclude that the "Omicron variant" of SARS-CoV-2 enters "BHK" cells indistinguishably from VSV-S.

      Thank you for pointing this out. We have rephrased the conclusion as “…omicron variant of SARS-CoV-2 enters BHK cells indistinguishably to VSV-S.”

      (g) Line 278. What happens to virus binding if the authors ectopically express hACE2 in CHO-K1 WT and CHO-pgsA-745 cells?

      We did not perform this experiment (see also our response to major comment 3 above).

      (h) Lines 280-281 and elsewhere (line 635). The authors state "pixantrone (PIX), a drug under clinical trial that binds HS to inhibit HS binding with proteins...." The authors should clarify that the drug is under clinical evaluation for cancer treatment because of its DNA intercalating activity (and not its HS binding activity) and cite any relevant ongoing trials. Also, in line 635, is reference #46 correct?

      As suggested, we modified this sentence as “pixantrone (PIX), a drug under clinical trial for cancer treatment due to its DNA intercalating activity, which can bind HS to inhibit HS binding with proteins”

      (i) Line 281-282. The authors should confirm in a Supplementary Figure that the anti-hACE2 antibody used blocks SARS-CoV-2-JN.1 binding to ACE2.

      In Figure 7D, we showed that PIX and anti-hACE2 antibody block SARS-CoV-2-JN.1 infection, suggesting that anti-hACE2 blocks SARS-CoV-2-JN.1 binding with hACE2.

      (j) Figure 7B. Can hACE2 co-localization be added to this panel?

      We did not perform this experiment. We addressed the role of ACE2 in these airway cells in subsequent panels of Fig. 7.

      (k) Figure 7C. The quantitative data show a 50% reduction in binding signal with pixantrone, whereas the microscopy images appear to show a much greater effect. Can more representative images be shown so that the data better corresponds?

      A ~50% effect is not as visually obvious as the current Fig. 7C. Therefore, we chose not to change the images. However, the statistics in Fig. 7C (right) clearly indicate an average effect of about 50%, as the reviewer pointed out.

      (l) In the Discussion, it is not necessary to use Figure callouts (as done in the Results). Please remove, with the exception of reference to the model.

      We prefer to call out Figures in the Discussion so that we can remind the readers where to find the data. The readers may choose to neglect these callouts. But some readers may read most the discussion part without going through the results carefully. In this case, the figure callouts may help these readers.

      (m) Please delete all references to "new" or "novel" models. It is unnecessary.

      As the reviewer suggested, we deleted “new” and “novel” throughout the revised manuscript.

      (n) Figure legends. Please make sure each panel indicates the # of independent experiments performed. This is included for some but not all panels. Also, a few panels use an unpaired t-test where an ANOVA with multiple comparisons is required (e.g., Figure 1G and S1).

      We agree that, for the three-group sub-comparisons shown within Fig. 1G and Fig. S1, the relevant analyses should account for multiple comparisons. In the revised manuscript, we therefore analyzed these predefined three-group subsets using ordinary one-way ANOVA followed by Dunnett’s multiple-comparisons test, with BHK or Vero used as the reference group as appropriate. The two-group comparisons were analyzed using unpaired two-tailed t-tests.

      Reviewer #2 (Recommendations for the authors):

      (1) It is well established that ACE2 is the receptor for SARS-CoV-2. The authors should not downplay this by saying it is "widely assumed", "typically thought", etc. The specific molecular details at various stages of entry (i.e, the role of HS) remain a bit unclear, but it is disingenuous to imply ACE2 is not the bona fide receptor by any conventional definition.

      The present work does not challenge the well-established view that ACE2 is the receptor for SARS-CoV-2 entry/infection, but suggests that HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, whereas ACE2 acts downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined. We made this point clearer throughout the revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docked at the HS clusters.

      As the reviewer suggested, we removed “assumed” and “typical” and clarify that our findings do not challenge this concept. For example, we modified the abstract

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely assumed as the receptor.”

      as

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely considered the receptor for cell-surface attachment and subsequent cell entry.”

      (2) When the authors state pseudovirus internalization is independent of ACE2, they should clarify that this is the case in cells not expressing TMPRSS2. Most physiologically relevant cell types express TMPRSS2, which will facilitate entry at the plasma membrane.

      As the reviewer suggested, we included the following sentence in the Discussion section: “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Line 130: "Endocytic internalization is the main viral infection pathway" and Line 180-181 is not precise and should be rephrased to include the cell types described in the figure. This may be true in BHK-ACE2 cells, but the evidence in this section does not show that this is universally or broadly true.

      We agree and have revised these sentences to limit the conclusions to the experimental context directly supported by our data. Specifically, our results support endocytic uptake as the major route leading to pseudovirus genome expression in the pseudovirus assays and cell types examined here, rather than as a universal entry mechanism for SARS-CoV-2 across cell types. We have therefore modified the subsection title and the relevant sentence in the Results to explicitly refer to the tested cells/assays.

      Across the revised manuscript, we have accordingly revised the text to distinguish initial virion docking/attachment from productive entry, to acknowledge ACE2 as the established receptor for productive infection, and to limit our mechanistic conclusions to the cellular systems directly tested here.

      (4) All bar plots should show individual dots (i.e., Figure 1G) to better reveal the variance of each dataset.

      While we respect the reviewer’s suggestion, this is not required in the journal style. We prefer plotting bar graphs without individual data points, which often makes it difficult to see the mean values.

      (5) Line 57: This is not accurate. HIV uses a receptor and a co-receptor, for instance.

      We thank the reviewer for noting this inaccuracy. We agree that viral entry frequently involves coordinated engagement of multiple host factors rather than a single receptor, for example, HIV requires both a primary receptor and a co-receptor. We have revised the statement in the Introduction (Line 57–58) to reflect that entry can involve receptors together with co-receptors and/or attachment factors, which collectively facilitate membrane fusion or endocytic uptake.

      In the Introduction (Line 57), we replaced the sentence with “Viral entry is often initiated by engagement of host receptors and associated co-factors that together facilitate subsequent viral membrane penetration.”

      (6) Line 60: "most" --> "many"

      As suggested, we have changed “most” to “many”.

      (7) Remove "clinically relevant" in reference JN.1, as JN.1 is not circulating currently. A more appropriate term could be "full-length" or "authentic", or "wild-type".

      As suggested, we changed it to “authentic”.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors omit to mention the work of Zhang et al, 2023 Nature Comms. "Host heparan sulfate promotes ACE2 super-cluster assembly and enhances SARS-CoV-2-associated syncytium formation". These authors also use PIXN and MTN compounds and define different mechanisms based on ACE2 clustering for virus entry. The authors should mention this work in the Discussion and try to reconcile the different findings.

      As suggested, we include the following discussion in the revised manuscript.

      “Consistent with this possibility, HS may promote spike-dependent ACE2 super-cluster assembly at the cell surface and enhance SARS-CoV-2–associated syncytium formation, suggesting that HS may organize ACE2 nanoscale architecture in a cell–cell fusion context [43].”

      (2) The authors should strengthen their case for the validity of HS-virion interactions as a therapeutic target by mentioning studies showing effectiveness of interference with HS-Covid interactions by heparin and other investigational drugs eg. first study to demonstrate heparin inhibition of SARS CoV2 attachment, Mycroft-West et al, Thromb Haemostatis, 2020; recent report of successful clinical trials of nebulized heparin, The Lancet, Sept 2025; and the superior efficacy of HS mimetic Pixatimod/PG545 compared to heparin (Guimond et al 2022 ACS Chemical Sciences).

      We thank the reviewer for this suggestion and add the following paragraph with citations the reviewer mentioned in the Discussion section.

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding.”

      (3) Figure 1a: incorrect label for antibody.

      Corrected, thank you.

      (4) Some misspellings noted in the manuscript, e.g., MINFLLUX, so please recheck the manuscript for typos.

      We have rechecked the manuscript and corrected the typos.

    1. eLife Assessment

      This important study investigates the mechanisms by which Mycobacterium tuberculosis suppresses protective Th17 differentiation during infection. Using ESX-1- and PDIM-deficient mutants of M. tuberculosis, which lack functional eccC1 and fadD28 respectively, the authors demonstrate that these virulence factors actively restrict IL-17 responses via a T-bet-dependent pathway, independent of overall bacterial attenuation or infection duration. While the precise molecular interactions by which ESX-1 and PDIM modulate Th17 differentiation warrant further elucidation, the experiments are rigorously designed, the findings are clearly presented, and the evidence is solid for a specific role for these factors in limiting protective IL-17 immunity in M. tuberculosis infection.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript examines the factors that restrict the induction of IL-17-producing T cells during Mycobacterium tuberculosis (Mtb) infection. The authors show that neither infectious route, nor duration of infection are responsible. But they do show that mice that lack the Th1-defining transcription factor, a finding consistent with prior reports in the field of immunology. They also show that 2 highly attenuated Mtb mutants in ESX-1 and PDIM, two well-known Mtb virulence factors, do induce IL-17 producing T cells. In contrast, Mtb mutants in mmpl4 are also similarly attenuated, but do not induce IL-17-producing T cells, suggesting that this property is not simply a result of attenuation but due to specific properties of ESX-1 and PDIM-deficient mutants.

      Strengths:

      (1) It is interesting that mice infected with ESX-1 and PDIM mutants have increased induction of Th17 cells.

      (2) Data is solid and convincing throughout.

      Weaknesses:

      There are two main criticisms:

      (1) B6 mice, compared to humans are known to be very Th1 skewed and the Th1 transcription factor T-bet is known to be a strong inhibitor of Th17 responses. Thus, these Th17 inhibitory factors may be stronger in B6 mice than humans, as many humans do make Th17 responses to Mtb infection.

      (2) The molecular insights about how Th17 induction is somewhat limited. Tbet induction is known to restrict Th17 development and this is a t cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. It is not clear these factors related to each other in restricting Th17 induction.

      Additional points:

      (1) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." One consideration, however, is that ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but have different CFUs in the lung-draining LN where T cell priming occurs? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction. Thus, without examining the LN, some questions remain regarding the conclusion that the altered Th17 response in the attenuated strains is not due to the attenuation itself. However, I agree the CFU in the LN probably reflects that in the lung, and if so, the author's conclusions would be sound.

      (2) Do LN cDC1 and high levels of IL-12 p35 manifest in mice infected with the mmpl4 mutant? Likewise do LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb). If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than to those infected with ESX-1 and PDIM mutants. These experiments would help provide more convincing evidence that the identified mechanisms are due specifically to outcomes regarding Th17 induction. However, I agree the author's conclusions are the most likely explanation given the current data.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors tackle an important question of why IL-17 production and TH17 responses are lower than expected during Mtb infection. The authors identify an axis of cross-regulation between TH1 and TH17 cells and provide data to support roles for Mtb virulence factors ESX1 and PDIM in promoting TH1 responses and/or suppressing TH17 responses.

      Strengths:

      The strengths include the significance of the work, the combination of host and Mtb genetic models to dissect the mechanistic basis for regulation of IL-17 production from T cells during infection, and the rigor of the experiments. There are a number of exciting findings from the work, including the cross talk between T cell responses and the impact of ESX1 and PDIM on these responses. It is particularly striking that that IL17a deficient mice partially rescue the attenuation of ESX-1 and PDIM mutants.

      Comments on revised version.

      The revised manuscript has tempered a lot of the language in the original text to more accurately state (and not overstate) interpretations of the data. The claim that the effect is independent of route of infection seems a little too large of a claim when only two routes were tested (aerosol and intranasal). And although the authors revised the results section to acknowledge the contribution of an IFNg-dependent suppression of IL-17 production from T cells, the abstract has not been updated and still claims that all effects are independent of IFNg.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Zilinskas et al seeks to understand the mechanisms underlying the ability of Mtb to suppress Th17 differentiation. As Th17 responses are needed for protective immunity against TB, this is an important topic of investigation. They use Mtb mutants that lack eccC1 (from ESX-1 locus) and fadD28 (encoding PDIM) and implicate a Tbet-dependent pathway by which Mtb modulates Th17 differentiation. The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is, however, unclear, which limits the novelty of the results.

      Strengths:

      Understanding how Mtb limits Th17 differentiation has implications for vaccine development. Comparative study of KO mice and Mtb mutants is a strength.

      Weaknesses:

      (1) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may reflect a difference in timing, since only 3-week data are shown; earlier and later time points would provide a better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, neutrophil data are important for clarifying the mechanism underlying the CFU outcomes.

      (2) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Fig 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6 etc) in the Tbet KO experiments.

      (3) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Fig 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      There are two main criticisms:

      (1) It is not clear how much the factors uncovered here are true beyond B6 mice. B6 mice, compared to humans, are known to be very Th1-skewed, and Tbet is a strong inhibitor of Th17-specific T cells. Many people make IL-17-producing T cells in response to Mtb infection.

      We appreciate the point that not all findings in mice are directly translatable to humans. The B6 mouse is widely used as a model organism for tuberculosis due to its tractability and the wealth of genetic tools available for this strain. While it is true that many individuals do produce Th17 cells after infection with Mtb, humans are still very Th1-dominant, and not all infected individuals produce Th17 cells. We can speculate that the mechanisms outlined in this paper may contribute to the reasons that Th17 responses are not more robust in humans, a finding that may be useful in guiding vaccine design in the future.

      (2) Very few novel insights are mechanistically revealed about how Th17 induction is restricted by Mtb. Tbet induction is known to restrict Th17 development, and this is a T-cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. How, if at all, are these factors related to each other in restricting Th17 induction? Also, the conclusion that it is not a result of attenuation is not completely convincing.

      While it is established that Th1 differentiation can inhibit Th17 differentiation, we believe that rigorously demonstrating this genetically in the context of Mtb infection remains important. Moreover, it cannot be assumed that IL-17 elicited by dampening the Th1 response can lead to enhanced control of infection. We view addressing this as a significant contribution. Furthermore, we also show that the ESX-1 and PDIM virulence factors are functionally linked by suppression of IL-17 responses. The effect is unlikely to be simply due to attenuation of the strains as an equally attenuated control strain does not elicit Th17 cells. We believe that these insights are both novel and important for understanding immune responses to Mtb.

      Other points:

      (1) The authors show that mice infected with a deficiency in ESX-1 have more IL-17-producing CD4 T cells in response to stimulation with an ESAT-6 peptide pool (Figure 3B). Because ESAT-6 is encoded by ESX-1, why do mice infected with this Mtb mutant have any ESAT-6-specific T cells? Is it an incomplete knockdown?

      The ESX-1 knock-out M. tuberculosis Erdman strain is a ΔEccC1 mutant. This strain can produce Esat-6 but cannot secrete Esat-6 out of the bacterial cell. Thus Esat-6 protein is present and able to be processed for MHC-II presentation. We also use Ag85b peptide pool stimulation and report similar effects as Esat-6 peptide pool stimulation.

      (2) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." I don't think the last sentence is necessarily true. I can imagine a scenario in which the induction of the Th17s is, in fact, due to the attenuation, and the Th17 induction still contributes to protection.

      We tested another attenuated M. tuberculosis strain with no known relationship with ESX-1 or PDIM, ΔMmpL4. This attenuated mutant fails to induce IL-17A–producing CD4 T cells to the same extent as observed in mice infected with ESX-1-deficient or PDIM-deficient strains, which is strong evidence that simple attenuation of virulence does not result in higher numbers of Th17 cells being elicited.

      (3) ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but what about the LN? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction.

      We acknowledge that this is a formal possibility, however we maintain that the phenotype is specific to ESX and PDIM mutants, rather than MmpL4 mutants. Even if this phenotype arises from a tissue-specific attenuation of ESX/PDIM mutants, it remains a specific phenotype of these mutants, and not all attenuated mutants, albeit less directly. More importantly, the observation that these mutants induce higher levels of the Th17-polarizing cytokine IL-23 from infected cells ex vivo suggests that this is not an indirect phenomenon.

      (4) Do LN cDC1 and high levels of IL-12 p35 in mice infected with the mmpl4 mutant? Likewise, LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb)? If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than those infected with ESX-1 and PDIM mutants.

      Because ΔMmpL4 and complemented strains resulted in T cell profiles that were not different from the wild-type, we did not measure mediastinal lymph node dendritic cell expression of IL-12 p35 and IL-23 p19 in infections with these mutants.

      (5) ESX-1 and PDIM are very different virulence factors - a protein secretory pathway and cell wall lipid, respectively? Mechanistically, how would mutants in these pathways give very similar outcomes regarding Th17 cells unless it was simply as an aspect of their attenuation? Perhaps, mmpl4 mutants simply differ in some aspects of their attenuation, such as bacterial burdens in LNs, or their interaction with cDCs?

      We are not the first to link phenotypes of ESX-1 and PDIM. Both systems have both been shown to be important for M. tuberculosis permeabilization of the host cell phagosome after phagocytosis, and for suppression of type I IFN responses, among other responses. Thus, these seemingly different virulence factors clearly work together to support specific virulence traits during infection. The exact mechanism of how ESX-1 and PDIM interact is not completely understood and is an area for future investigation.

      Reviewer #2 (Public review):

      The following conclusions and interpretations should be revisited, rephrased, and re-evaluated:

      (1) The manuscript neglects to analyze T cell responses in the dLN, which is the critical site where these responses are initiated (only DC cytokine production is measured in the dLN). The differences in the lungs could reflect trafficking of T cells to the lungs, local lung T cell responses, or durability of the T cell responses in the lungs. The authors state in the last results section that "These results indicate that the ESX-1 and PDIM virulence factors impact naïve T cell differentiation at the draining mediastinal lymph node..." but T cell responses are never measured in the dLN.

      Due to the limited size of the mediastinal lymph node at 3 weeks post infection, we were unable to obtain enough cells for both myeloid cell analysis and T cell analysis, as we perform staining for these panels separately due to the decrease in viability of myeloid cells observed during T cell restimulation. In addition, because T cells in the lung are the population of cells most critical for mediating the outcome of infection, we believe analyzing the T cell response in the lymph nodes though interesting, is not crucial for this study. We have edited the manuscript to be clearer, as suggested by the reviewer.

      (2) Figure 2: The authors state that "Importantly, IFN-γ deficient mice did not exhibit elevated levels of IL-17A producing CD4 T cells demonstrating that IFN-γ production is not the mechanism by which Th1 T cells limit a Th17 response during Mtb infection", but the difference is significantly different and even more obvious in Panel B. In fact, if the Panel D y-axis was on a log scale, the Ifng-/- would likely look more like Tbet-/- than WT. Based on this data, it seems like IFNg is having an effect and should not be completely discounted. Does the deletion of Ifng affect the number of Tbet+ T cells?

      We agree that the IFN-γ<sup>-/-</sup> have only 5x more IL-17 producing CD4 T cells than WT mice while Tbet<sup>-/-</sup>mice exhibit a 25-fold increase compared to WT. We have added this information to the text, and now point out that IFN-γ production is not the sole mechanism by which Th1 T cells limit a Th17 response during Mtb infection.

      In addition, the deletion of Tbet results in an increased number of IFNg+IL-17+ double positive T cells (Figure 2B), in addition to a sizable IFNg single positive T cell population maintained in the Tbet-/- mice (10x the negative control of Ifng-/-). Is this why Tbet deletion is not as severe as Ifng deletion, because T cells are still making IFNg?

      It is possible that the residual IFN-γ produced by T-bet-deficient animals contributes to their relatively modest susceptibility to infection. However, our data show that deletion of IL-17 in this background renders T-bet–deficient mice nearly as susceptible as IFN-γ deficient mice, arguing that the remaining IFN-γ is not a major protective factor.

      Along these lines, the statement in the text that, "Tbet-/-Il17a-/- mice completely lacked both IFN-γ producing...." T cells is not supported by the data in Figure 2C. Tbet-/-Il17a-/- mice look to have more gamma-producing T cells than Tbet-/- mice (which is already 10x the negative control of Ifng-/- in panel 2B if one includes the gamma single positive and IFNg/IL-17 double positive).

      We have amended the language in the text to be more consistent with the data.

      (3) In the Results sections describing Figures 3, 4, and 5, the authors equate IL-17 production by T cells with TH17 responses and IFNg expression with TH1, but Tbet and RORgt expression in the T cells should be measured to make conclusions about TH1 and TH17. Or the authors can rephrase their findings to specifically state the observations as IFNg or IL-17 expressing CD4+ T cells.

      We believe that calling a CD4 T cell in the lung that is producing IFN-γ (and not IL-17) a Th1 cell is appropriate. Potentially confounding cells include those which also produce IL17, which we have ruled out, or T<sub>FH</sub> cells that may be common in lymph nodes but are not common in lungs at this time point and under these conditions.

      (4) Conceptually, do the authors think that ESX1/PDIM promotes TH1 responses and this blocks TH17 or are ESX1/PDIM blocking TH17 responses directly, allowing for increased TH1 responses? It would be helpful to clarify the model in this regard, describe how the data supports one model or the other, and then make sure the language is consistent throughout. Can these effects on T cell responses be tested and recapitulated in vitro using infected APC and T cell co-cultures?

      While it is possible that PDIM and ESAT-6 suppress Th17 through promotion of Th1 differentiation, we do not have data to support this model currently. However, we have added a comment making this point to the discussion.

      Reviewer #3 (Public review):

      Weaknesses:

      (1) The authors should acknowledge and reference key findings from the literature that have identified suppression of Th17 differentiation as an Mtb virulence mechanism, e.g., the role of the Hip1 protease and CD40 signaling (Madan-Lala JI 2014, Sia Plos Path 2017, Enriquez iScience 2022) and Khader JI 2005, showing the requirement of IL-23 for Th17 responses in vivo in a TB mouse model.

      We thank the reviewer for pointing these references out and have added them to the discussion section of the manuscript.

      (2) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may be due to a difference in timing, since only 3-week data are shown; earlier and later time points would provide better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, data on neutrophils are important for clarifying the mechanism for the CFU outcomes.

      We agree that it is surprising that, in the context of TB, Th17 responses are protective whereas excessive neutrophil recruitment is detrimental to the host. In IFN-γ–deficient mice, neutrophils are recruited and contribute to the increased susceptibility of this strain (PMID: 21967766). In separate work from our lab, we have shown that the phenotype of neutrophils recruited to the lungs during Mtb infection influences disease outcome (PMID: 40937719). It is possible that differences in the host environment and the timing of the response shape the effects of neutrophils on the host; these and the other questions raised by the reviewer will be the subject of future studies.

      (3) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Figure 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6, etc.) in the Tbet KO experiments.

      While these are interesting points, investigating mechanisms of Tbet-dependent suppression of IL-17 is beyond the scope of this study.

      (4) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Figure 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants. Additionally, what was the level of Type I IFNs elicited by these mutants?

      We included the MmpL4 knockout Mtb Erdman strain as a control to ensure that attenuation of mutants is not the cause of the increase in IL-17. We also showed that eliminating type I IFN signaling by deleting its receptor has minimal impact on Th17 differentiation, even in the context of a host that produces excess type I IFN. Therefore we do not believe that type I IFN elicited by these mutants is explanatory for the phenotype.

      (5) Since macrophages have been implicated in the reduced cytokines seen in the ESX-1 mutant, IL-23 and other cytokine data on lung macrophages would complement the DC data.

      Because dendritic cells are primarily responsible for priming CD4 T cell responses, we believe that this result in macrophages would not substantially alter our conclusions. That said, it was demonstrated previously that macrophages infected with ESX-1 mutants produce less IL-12p40, a subunit of IL-23.

      (6) Figure 5. There are many fewer DCs overall in the eccC1 and fadD28 mutant groups, which could account for the increased % IL-23p19 in DCs (5D). What were the levels of IL-23 in DC1s?

      The amount of IL-23 p19+ in type I conventional dendritic cells (cDC1s) was near zero as shown in supplementary figure 6A. cDC1s are known to not express IL-23 p19 in mice.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What do the authors mean by the alternative secretion part of "ESX1 Type VII alternative secretion system" that they refer to?

      Bacterial alternative secretion systems facilitate the export of proteins from the bacterial cell independent of the canonical Sec-dependent secretion system required for export of most secreted bacterial proteins across the inner membrane. However to avoid confusion, we have removed the word alternative.

      (2) Not sure naïve fits in this sentence at the end of the introduction: "Furthermore, we observe a strong Th17 response during infection with ΔESX-1 or PDIM lacking Mtb in naïve mice....".

      We have removed the word Naïve.

      (3) Figure legend for 1A-C says analysis performed at 21 dpi, but the figure shows the time course.

      We have corrected this error.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1 should show the non-stimulated flow plot.

      We have added the unstimulated samples.

      (2) The % IL-17 in the flow plots is not consistent across Figures 1, 2, and 3. Not sure why the scales for the Y-axis for IL-17 differ so much between Figures 1 and 2/3. IS there a technical issue with compensation?

      We did not experience any difficulties with compensations. These experiments were done over several years of work. For every experiment, new single-color controls were used and gating was done with the FMO gating strategy. Minor variation such as we see here is not surprising.

      (3) Discuss Yeh et al J Neuroimmunol 2014- show that IFNγ inhibits Th17 differentiation and function via Tbet-dependent and Tbet-independent mechanisms.

      We have added this reference to the manuscript.

    1. eLife Assessment

      This useful study reports findings that support the use of the Open Field Test in Drosophila as a model to study "emotion-like states", which are behavioral responses to several stressful or aversive treatments, and resilience upon their subsequent removal. Behavioral data, by employing established stress-causing treatments and genetic manipulations, are solid. The main advance of this work over previous Drosophila work using a similar experimental setup is systematic analysis of various stressors.

    2. Reviewer #1 (Public review):

      Summary:

      Animal behavior is continuously influenced by the internal state moment by moment, including emotion primitives as the authors pointed out. Although emotion is a more human-related state, evolutional conservation is undeniable, which can be inferred by the behavioral manifestation. To further elaborate the neuronal mechanisms of emotion primitives, the simplest behavioral parameter related to emotional primitives should be well characterized. In this study, the authors described in detail of wall-following behavior (WAFO) and the total walking distance (TOWA) using flies after subjecting them to various conditions or flies being genetically manipulated according to the previous reports that could affect emotion primitives. Overall, the study is well designed and structured. In addition, the discussion on emotion primitives will be of value to the field.

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      Conceptual concerns:

      My suggestion to strengthen the authors' conclusion that "TOWA can be interpreted as a behavioral proxy for exogenously induced arousal" was to show that an increase in TOWA after stress exposure can be observed in a small (1-cm) arena that acts as an exogenous arousing stimulus, but not in a larger arena (>6.6 cm) where such arousing effects are absent. This comparison would demonstrate that basal locomotor activity measured in larger arenas is not altered by stress, whereas the additional component observed in smaller arenas reflects stress-induced internal state. Therefore, the authors would be able to distinguish clearly the effects of stressors or experiences on either simple locomotion or an emotion-like internal state. Then the future works can follow this protocol using smaller and larger arena to assess emotion-like internal state.

      I appreciate the significant authors' efforts to monitor TOWA using arenas with different diameters up to 6 cm. However, the conclusion was unfortunately the same as that obtained using the 1-cm arena. As the authors commented, flies do not show persistent and quantifiable wall-following in arenas larger than 5.8 cm, which limits further examination of this question. I personally agree with the authors' interpretation, but I hope that the authors obtain more definitive experimental contrasts to support this claim in the future study.

    3. Reviewer #2 (Public review):

      Summary:

      In terms of data, the revised manuscript is by and large the same, though the authors added new experiments examining the effects of the Open Field Test (OFT) arena size (Fig. 1 Supplement 1) and sex and mating status (Fig. 6). The authors have also provided textual revisions, partially addressing my previous major criticism about novelty over the work of Mohammad et al., 2016, Curr Biol. They argue that the main advance is the systematic inclusion of Total Walking (TOWA) data (e.g. Introduction, page 6 top in the tracked changes document). While I am still not convinced that the findings represent a huge leap forward over that previous work, the authors' systematic analysis is very nice and may prove of use to those seeking to develop Drosophila as a model for studying emotion primitives.

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stress-like emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      Weaknesses:

      The authors have addressed some of my previous points with textual revisions and in their rebuttal. As stated above, the conceptual advance over Mohammad et al. remains in my opinion limited, but I appreciate that this point is now clearly discussed in the manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      We appreciate this assessment and fully agree that the simplicity of the assay and the demonstration of its scalability and reproducibility are a strength.

      Weaknesses:

      The weakness of the study is the lack of further experiments to support their assumption related to TOWA. The authors suggested that TOWA can be interpreted as a behavioral proxy for exogenously induced arousal. However, it could be interpreted as higher activity, although the authors argued that the circadian clock increasing locomotor activity around ZT0 and ZT12 does not affect TOWA, and therefore TOWA is not related to the locomotor activity per se. As the author cited, flies lose locomotor activity in the circular arena of 6.6 cm in diameter, whereas they continuously move during a 1-h recording in the authors' arena of 1 cm in diameter.

      I would agree that the arena of 1 cm in diameter, but not 6.6 cm in diameter, serves as an exogenous stimulus inducing arousal, and TOWA is manifested by arousal. However, TOWA would also be affected by other behavioral parameters, including the activity, motivation for exploration, or perception of the space. Therefore, it could be reasonable to re-examine some of the flies tested in this study in the circular arena of 6.6 cm in diameter. If arousal is biased by the components presented in Figure 6 and TOWA can assess mainly exogenously induced arousal, the treatment altering TOWA in the arena of 1 cm in diameter would not affect their behavior in the arena of 6.6 cm in diameter. My concern is that Figure 6 may demonstrate too simplistic a diagram to interpret the results. I would suggest adding the experiments using the arena of 6.6 cm diameter or softening the argument.

      We are grateful that you prompted us to investigate the relation between TOWA and arousal and different arena diameters in more depth. Based on your comments, we compared naïve and stressed behaviour between arena diameters of 1, 2.2 and 5.8 cm. The sizes were chosen as to optimally comply with the camera field of view in our setup. Naïve flies showed stable locomotor activity throughout the 60 min of recordings in the different arenas (new Figure 1 – figure supplement 1 A-B’’). Moreover, no significant difference in TOWA over the first 10 min was found between the different arena sizes (new Figure 1 – figure supplement 1 C). This suggests to us that arenas with a diameter smaller than 6.0 cm (and not only with a 1 cm diameter) induce some form of activity that resembles stimulated activity as defined by Meehan and Wilson (1987). A mechanical shock (shake) resulted in significantly increased WAFO in all arena sizes (new Figure 1 – figure supplement 1 D-D’’). We did not test the effects in larger arenas > 5.8 cm, as we found that flies do not longer show persistent and quantifiable wall following. We started to see that also in few flies in the 5.8 cm arena – these few flies were excluded from our analysis presented in new Figure 1 – figure supplement 1. Unlike the naïve response, the stress-induced response in WAFO and TOWA appears to be transient (new Figure 1 – figure supplement 2), which is in line with the definition of emotions as a transient state.

      Reviewer #2 (Public review):

      Summary:

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stresslike emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      We are glad about this assessment and share your opinion.

      Weaknesses:

      The conceptual advance of this research is unclear, as previous work (Mohammad et al., 2016, Curr Biol.) carried out similar treatments and manipulations and reached largely similar conclusions.

      Thank you very much for bringing this up. We rewrote respective parts of the introduction (second last paragraph) and discussion to more clearly outline the advances over the previous work by Mohammad et al. 2016. While our study builds upon Mohammad et al. 2016, the conceptual advance and novelty is that we constitute and treat TOWA as a second and independent dimension equal to WAFO in the OFT. Mohammad et al. had measured locomotor activity (reported as average speed in their paper (total distance walked/time of recording), but primarily to test the dependency of WAFO on locomotor activity. They found that WAFO metrics were poorly correlated with average walking speed, showing a significant degree of independence of both measures – a finding that our results confirm. However, unlike us, they did not consider average speed/TOWA further for their analysis, possibly because they focused on anxiety-like behaviour while our study looked broader on emotion-like behaviour in general. We further used a round (not square) arena to exclude “cornering” in order to reduce the complexity of the assay, which may explain differences of observed speed/TOWA between our studies.

      Moreover, while WAFO is a good proxy for 'stress', I am not convinced that TOWA necessarily represents an emotional state in all cases. Indeed, as the authors themselves acknowledge, changes in total walking may be associated with other factors, such as starvation-induced hyperactivity, physical exhaustion after sleep deprivation, increased sex drive after mating, alcohol sedation, etc.

      Your comment raises a question in comparative research on emotions which is very difficult if not impossible to conclusively answer. At first sight, the most conservative stance seems to be to completely disregard the idea of emotions and affective experiences in animals. This, however, would mean that we cannot use animal models to study the basics of emotions (= emotion primitives) and would ignore that by all likelihood emotions are a product of evolution and hence should exist at least in more basic forms in animals. Obviously, we have no means to ask flies or any other animal whether they connect “hunger” or “mating” to a feeling or an emotional state (which must not be conscious) but can only observe the behaviour. We here adopt the often-cited “Pankseppian” view (based on the book of Jaak Panksepp: “Affective Neuroscience”) and firmly believe that – in order to fully understand how the brain drives behaviour- we also need to take affective states into account that bias behaviour towards adaptive responses.

      In short, we are unfortunately unable to give a clear and definite answer to your comment whether TOWA represents an emotional state in all cases. Perhaps you are right. We believe, however, that a “Pankseppian” view is adequate, and we may ask what evidence exists that shows that starvation-induced hyperactivity or post-mating is not associated with an affective emotion-like state in the fly or any other animal.

      Another unclear point is the interpretation of some unexpected results, such as the finding that both serotonin transporter overexpression and its knockdown give the same phenotype.

      Thank you very much for this comment, which we also received by reviewer #1. As suggested by the other reviewer, a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT may be a differential effect on the different serotonin receptor subtypes expressed in the brain. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – both higher and lower than optimum levels result in behavioural impairments in adults (see e.g., Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Finally, there are some issues with the use of the OFT in rodent research (e.g., inconsistent effects of anxiolytic drugs; see Rosso et al., 2022, Neurosci Biobehav Rev., for a meta-analysis). These should be explained to place the Drosophila findings in their appropriate context.

      Thank you very much for bringing this systematic review to our attention which assessed the usefulness of various behavioural tests including the OFT to study the effect of anxiolytic drugs in rodents. Overall, the review casts “serious doubt on both construct and predictive validity” of behavioural tests for anxiolytics. While diazepam (the only drug used in our study) was the drug with the most consistent effects across the analysed behavioural assays, only 59% of the OFTs revealed significant effects. We were already aware of earlier findings in the same direction (Prut and Belzung 2003 10.1016/s0014-2999(03)01272-x), but as we only used one drug did not include a discussion in the manuscript. Unfortunately, the number of studies employing the OFT in flies is very small and does not yet allow for a similar comparison. We now changed the respective sentences in the discussion:

      “In rodents, diazepam mostly but not consistently leads to an anxiolytic response in the OFT behaviour which questions the usefulness of the OFT for testing anxiolytic drugs (see (Prut and Belzung 2003; Rosso et al. 2022)).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Overexpression of SerT suppressed increased WAFO after electric shocks. However, knockdown of SerT did similar. It is worth reporting this finding, but possible mechanisms could be mentioned in the manuscript. For example, given that flies carry five serotonin receptor genes, upregulation of overall serotonin level may affect the specific serotonin receptor, but downregulation of it may affect other receptors, leading to unexpected outcomes. Exploring any other possibilities could be better to add to guide future research.

      Thank you very much for this comment and the suggestion of a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – higher or lower levels result in behavioural impairments in adults (see e.g. Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Reviewer #2 (Recommendations for the authors):

      (1) The advance over Mohammad et al., 2016, Curr Biol. must be clearly and emphatically articulated in the Introduction and/or Discussion. It is otherwise impossible to appreciate what is conceptually novel about this work.

      We have now rewritten parts of the two last paragraphs in the introduction to make the advances clearer. As outlined above, we conceptually advanced the analysis of OFT behaviour by integrating TOWA as a second and independent dimension in our analysis. During the revision process, we have spent great effort to better characterise the nature of the locomotor activity encountered in the OFT (Figure 1 – supplementary figures 1 and 2). Also, this is an advancement over previous studies, including Mohammad et al. 2016. We further included new results on the general effect of neuropeptides (silver mutants, impaired in neuropeptide processing).

      (2) The most important metric for a stress-like emotional primitive is WAFO. Therefore, figures should be revised in a way that highlights WAFO differences. Figure 4 is a good example of this. In contrast, in Figures 1-3, the WAFO box-and-whisker plots are very small and obscured under the raw WAFO-TOWA plots, which are difficult to see (especially given the light blue background of all plots) and redundant. I strongly recommend just showing the WAFO box-and-whisker plots for the sake of visibility, clarity, and brevity.

      Although we understand your reasoning, we would like to stick with the old figures as we consider TOWA as important to characterise the OFT response as WAFO (see our comments above regarding the conceptual advances of our study over Mohammad et al. 2016). It is true that Figures 1-3 are small, but at least in our print-out well legible. Further, it appears that in the current version of the eLife system the resolution is downsampled. In addition, we anticipate that in the version of record figures can be enlarged online as in other eLife articles.

      (3) Given the criticisms against the OFT in rodent research, one of which is inconsistent effects of anxiolytic drugs (Rosso et al., 2022, Neurosci Biobehav Rev.), it may be useful to expand pharmacological treatments beyond diazepam.

      As our focus is not on the testing of anxiolytic drugs and since it was already very difficult to be granted access to diazepam (we are not at a medical institution), we refrained from testing further drugs. Moreover, as rightfully mentioned by you, the OFT may not be the best test for the efficacy of anxiolytic drugs. On the other hand, diazepam was the most consistent anxiolytic in the OFT in rodents (see Rosso et al. 2022).

      (4) The authors should show results of the effects of at least some stressors/punishments on WAFO/TOWA of female flies to understand if observed effects are sex-specific.

      Thank you very much for bringing this topic to our attention. To test whether the effects are sex-specific, we now performed several new experiments. First, we compared the naïve OFT response of mated and unmated males and females (see new Figure 6). This revealed that without prior stress treatment, the WAFO response is independent of sex and mating status. In contrast, the naïve TOWA response turned up to be sex- and mating state-specific (see new Figure 6). To test whether the mated females show a different stress-induced OFT response to males, we applied mechanical stress (shake) that we had also used to assess the effects of arena diameter (Figure 1 – supplementary Fig. 1 D-D’’). After a first round of shaking, females showed increased TOWA, but WAFO was unaffected. A second round of shaking, however, led to a significant increase in WAFO and TOWA. This suggests that the OFT response is qualitatively similar between the sexes and mating status, yet the threshold for elicited responses differs between males and females. We now added a respective paragraph to the main text in the results section plus a new figure (Fig. 6).

    1. eLife Assessment

      This important study addresses a timely issue at the intersection of mitochondrial and telomere biology by focusing on the relationship between naturally occurring variants in the mitochondrial genome and telomere length. This work thereby provides a conceptual and experimental framework for investigating communication between mitochondria and telomeres. Using an innovative transmitochondrial cybrid approach, the authors provide evidence that mitochondrial DNA variants influence telomere maintenance through effects on mitochondrial function, reactive oxygen species, and NAD⁺-dependent repair processes. The evidence supporting the central conclusion that mitochondrial genotype influences telomere-associated phenotypes is convincing and is strengthened by the use of complementary functional and rescue experiments. However, some of the mechanistic interpretations and broader conclusions regarding telomere length inheritance in humans would benefit from additional donors and longitudinal analyses following cybrid generation, or more cautious framing.

    2. Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result. Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes? In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

    4. Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R²=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R²=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      We appreciate the reviewer raising the important link between pathogenic POT1 variants and lymphoid malignancies driven by elongated telomeres. We would like to clarify, however, that our introduction focused not on the pathological telomere attrition characteristic of telomere biology disorders, but rather on the natural, non-pathological variations observed in individuals with baseline telomere lengths on the shorter end of the spectrum.

      Nevertheless, we will include this observation regarding patients with long telomeres in the introduction to underscore the complex relationship between telomere length and tumorigenesis.

      The two following publications by the Armanios lab will be added:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      Because this study investigated the potential maternal inheritance of TL via the mitochondrial genome, our cohort consisted predominantly of female donors (6/7). Donor #5, the husband of Donor #7, was the only male included. We will update Figure 1D to include the sex of each donor and clarify this rationale in the text.

      Notably, our analysis revealed minimal influence of sex on TL, which cannot account for the observed differences between the extreme groups.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      We will do so.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result.

      We thank the reviewer for this suggestion. Following their advice, we used our metaphase spread FISH analyses to quantify chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells. However, because Cybrid 2 metaphase spreads were of insufficient quality for adequate chromosome analysis, we will restrict our telomere fusion comments exclusively to Cybrid 1 and include the quantifications in our revised manuscript.

      Author response image 1.

      The number of fusions and total chromosomes analyzed are indicated on each bar.

      Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      While this is a strong argument, the telomeres within these cybrid models are critically short. Consequently, the FISH signal intensity falls below the threshold required for reliable quantification of telomere-free ends.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      The reviewer is correct that we have not formally demonstrated this in our experimental system. Instead, our assumption was based on the fact that the initial phase of Rho0 cell repopulation involves a temporary ROS burst, partly driven by incompletely assembled ETC supercomplexes that are known to elevate ROS levels (Maranzana et al, 2013). This is supported by our experimental observation that the NAC antioxidant, combined with NR, potently inhibits telomere shortening in cybrids with low CI activity. We propose to add this explanation in the revised manuscript.

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes?

      We cannot fully explain the discrepancy, but we indeed suspect in vivo oxidative stress differs significantly from cell cultures (21% O<sub>2</sub>). Additionally, early cybrid replenishment involves an unknown telomere elongation step that may offset mitochondrial ROS-induced shortening.

      To further investigate this question, we measured telomere length across varying population doublings (PDs) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include (as Supplementary figure) and discuss these data in the revised manuscript.

      Author response image 2.

      In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      We excluded Rho0 cells from our superoxide measurements because these cells were grown in a different culture medium. Instead, we focused on comparing mitochondrial ROS across cybrids to accurately correlate these values with the telomere length of the corresponding donors’ lymphocytes. Consequently, Rho0 cell measurements would not have contributed to this correlation analysis.

      We propose to discuss this in the revised manuscript.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

      We thank the reviewer for this constructive feedback and entirely agree that including more donors would have added depth to our findings. While we acknowledge this limitation, pursuing this further is currently impossible without acquiring new ethical approvals and establishing fresh collaborations with clinicians.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      We limited our TL measurements to the earliest viable time point after cybrid formation to avoid the confounding effects of cellular adaptation in culture. For instance, the early telomere shortening observed in Cybrid 1 and Cybrid 2 was later alleviated—likely due to the upregulation of the NAD+ salvage pathway genes NAMPT and NAPRT1.

      To accurately capture the effects specific to cybrid formation, we isolated 4 independent clones per donor, all of which showed highly consistent TL values, as shown in Figure 2A and S3D.

      While we acknowledge the reviewer's point about multi-passage longitudinal analyses, we did, in fact, measure telomere length across varying population doublings (PDs) in Cybrid 3 (high mito ROS levels), 6 and 7 (low mito ROS levels) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include a Supplementary figure and discuss these data in the revised manuscript.

      We further propose to carefully edit the manuscript so as to clarify the text and, whenever possible, add quantifications. Among others, as suggested by Reviewer #1, we will add the quantification of chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells in our revised manuscript (See Author response image 1).

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

      Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R<sup>2</sup>=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      We thank the reviewer for this insightful comment. As suggested, we will explicitly describe the sampling structure upon introducing the correlation analysis. Furthermore, we have conducted the requested leave-one-out analysis:

      - removing Cyb3: R<sup>2</sup>=0.730; p=0.0302

      - removing Cyb6: R<sup>2</sup>=0.782; p=0.0194

      - removing Cyb1: R<sup>2</sup>=0.740; p=0.0278

      While we agree that additional donors would enhance the study, further experiments are however currently impossible without new ethical clearances and additional clinical partnerships.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R<sup>2</sup>=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      We acknowledge that our study does not establish the in vivo role of CI activity in TL regulation. Our abstract specifically highlights an in vitro phenomenon: “Under the specific conditions of cybrid formation, which involve a metabolic shift from glycolysis to oxidative phosphorylation, mtDNA variants associated with reduced CI activity induced rapid telomere shortening, …”.

      We are nevertheless happy to revise the text to clearly separate our in vitro results from in vivo biology as requested.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

      We agree that the evidence for the AT6 A177T inheritance claim remains inconclusive. To clarify, we do not argue that this mitochondrial variant is solely responsible for longer telomeres; indeed, the mtDNA genome of donor #1 suggests otherwise. Furthermore, the phenotypic impact of such mtDNA variants likely depends on nuclear variants in other telomere-related genes (e.g., hTERT or hTR), meaning AT6 A177T may not consistently result in elongated telomeres. Unfortunately, our ethical protocol precludes screening additional unrelated donors. We will revise the text to soften our statements accordingly.

      New references:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      Maranzana E et al. Mitochondrial respiratory supercomplex association limits production of reactive oxygen species from complex I. 2013. Antioxid Redox Signal 19, 1469-1480.

    1. eLife Assessment

      This fundamental manuscript describes a key role for the integrated stress response-regulated transcription factor CHOP in regulating liver biology in response to endoplasmic reticulum stress through both the downregulation of transcription factors involved in regulating hepatic identity and altering the capacity for integrated stress response and unfolded protein response signaling to induce protective signaling. The data supporting this model is convincing, but including some additional discussion on the mechanism and importance of the work in the context of the published literature would be helpful to better define the complex importance of CHOP signaling. This work will be of interest to a wide range of biologists interested in liver biology, stress-responsive signaling, and ER stress.

    2. Reviewer #1 (Public review):

      Summary:

      The predominant view on CHOP's functions during ER stress is that it promotes cell death. This is in contrast to a handful of reports in the literature that claim that CHOP is a positive regulator of protein synthesis during chronic ER stress, and therefore is part of the adaptation program to ER stress. These previous studies were performed in tissue culture cells. Velarde and co-authors have used a mouse model of induction of mild ER stress to study the function of CHOP in hepatocytes.

      Major strengths and weaknesses of the methods and results:

      The authors use state-of-the-art mice to manipulate (i) CHOP and (ii) ATF6, a protective factor of ER proteostasis, and address the hepatocyte responses to mild ER stress in vivo and in cultures. Validated gene expression programs are well correlated to liver pathology in the mouse models. This is a very well-done study.

      The authors clearly show that CHOP transitions hepatocytes under mild ER stress to a chronic ISR state, which is phenocopied by ATF6-depleted hepatocytes. So the conclusion that CHOP exacerbates ER stress in hepatocytes during mild ER stress is correct. It is also clear that CHOP targets negatively the transcription of hepatocyte identity genes, which opens a new direction of studies on the function of CHOP in secretory cells in general.

      Conclusion:

      This is a significant study that will benefit different research fields, and specifically studies on proteostasis, as was recently highlighted in Nat. Str. Mol. Biol. by experts in the field.

      To this reviewer, the importance of the study is that it links the function of a transcription factor (CHOP) to stress intensity (mild versus severe) in a physiological experimental model (hepatocyte function and pathology).

    3. Reviewer #2 (Public review):

      The Unfolded protein response (UPR) and related integrated stress response (ISR) are critical signaling systems for cell survival in response to acute stresses. While the UPR directs critical adaptive gene expression, certain chronic stresses switch this pathway towards cell death and disease. An important question concerns the mechanisms by which the UPR switches from being adaptive to maladaptive. Prevailing models focus on the transcription factor CHOP (DDIT3 or GADD153), whose levels are enhanced via the UPR, and extended/amplified amounts of CHOP are suggested to boost death-related gene expression. However, the literature and this manuscript point out a number of observations that do not neatly fit with this model, suggesting that there are still unresolved processes by which CHOP adjusts cell outcomes via the UPR.

      This manuscript features a nice hepatocyte-targeted knockout of CHOP to discern the contribution of CHOP in the transition between adaptive and maladaptive outcomes. The key ideas presented in this study are that CHOP-directed gene expression is focused on protein synthesis, metabolism, and hepatocyte identity. In the progression of the UPR, CHOP expression can lead to resumption of protein synthesis, which can assist in the translation of the UPR-directed transcriptome, which includes ATF6/XBP1-directed genes that aid the processing capacity of the endoplasmic reticulum (ER). However, enhanced nascent protein can further stress the ER. CHOP directs gene expression in both the first phase- acute and second phase-chronic in the UPR, and the pivotal decision lies in the transition between the phases.

      Overall, the manuscript includes some new ideas as well as refinements of earlier ones for CHOP-determination of UPR-directed cell fate. The CHOP-hepatocyte knockout mouse model helps to delineate the different tissue functions of CHOP, which has been a problem for some earlier studies. The manuscript progression of experiments is solid, and experimental design and documentation are rigorous. The manuscript text is largely clear, but there are portions that would benefit from fuller explanations of ideas.

      There are three points of concern. First, the manuscript model (Figure 7) lays out a timeline for the progression of the UPR between two phases. The study is not always clear about the times assayed, and there appears to be a single time point for measurements. Second, there is emphasis on protein synthesis changes in the model. It is true that the literature argues that resumption of protein synthesis concurrent with stress damage (i.e., GADD34-directed gene expression) is a key reason for the potentially debilitating effects of CHOP (e.g., Marciniak et al 2004, Han et al 2013). However, the manuscript does not feature protein synthesis measurements. Inclusion of bulk protein synthesis measurements in the context of this model system would strengthen the study and support for the model. Finally, for this reviewer, some of the most interesting ideas center on CHOP-directed transcription of genes that regulate hepatocyte identity. There is solid evidence for direct CHOP regulation of these genes, but the manuscript does not really develop and test the ramifications of these networks on cell fate during ER stress.

      Reviewer Concerns:

      (1) The abstract packs in a lot of information. The ideas would not be clear to a general reader. Furthermore, the UPR and ISR are referred to in the second-to-last sentence, but not defined earlier in the abstract.

      (2) There are some typos/grammar concerns.

      (3) ATF4 diminished with CHOP-depletion (Figure S2A). What is the mechanism here? Does this complicate the analysis of CHOP-directed gene expression? How does this fit with Figure 6J? The timelines for TM treatment are critical. The authors should more fully explain the time courses in the experiments.

      (4) Figures 2 and 3: There is a discussion on enhanced protein synthesis with loss of CHOP (reduced GADD34 expression). What is the time point - 8 hours TM? Emphasize, explain, and justify time points of experiments here and in later panels. It would strengthen the model with direct measurements of protein synthesis. The authors could include GADD34 protein measurements in these panels. Figure 3 - panel D - some abbreviations are not standard.

      (5) Figure 4: One of the most interesting in the manuscript is the transcription factors downstream of CHOP that are linked with hepatocyte differentiation and metabolism. The manuscript would be bolstered by developing some of these target genes into the Figure 7 transition model.

      (6) Figure 6: The comparison of CHOP and ATF6 target genes is a highlight of the manuscript. The literature on this topic is complex, and there are some suggestions that CHOP can be downstream of ATF6. Furthermore, there were some earlier models by Walter and others about extended induction of Perk (death) vs induction of other UPR sensors (survival) (e.g. PMID: 17991856). It would be helpful in the Discussion to delineate between these models and their critical differences.

    4. Reviewer #3 (Public review):

      In this manuscript, the authors aim to understand the function of the transcription factor CHOP, which is known to promote cell death during severe stress in the ER. The authors note that CHOP is induced during less severe stress, but its functional output is not well understood in these cases. Here, they study the effects of conditional knockouts of CHOP in hepatocytes of mice challenged with chemical inducers of ER stress.

      Tunicamycin (an ER stress inducer) injection leads to the upregulation of CHOP and lipid accumulation in the liver, but no significant cell death in the experiments outlined here. Conditional knockout of CHOP results in a number of differences in the way hepatocytes respond to stress, notably resulting in lower steatosis.

      There are two main findings supported by the data presented here. First, the authors show that CHOP suppresses the expression of ONECUT, a master regulator of hepatocyte differentiation and metabolism, during ER stress. They show by ChIP-seq that CHOP binds to the promoter region of this gene, and by RNA-seq that ONECUT expression is suppressed by ER stress in a CHOP-dependent manner. Many predicted targets of ONECUT1 were also suppressed by ER stress in a CHOP-dependent manner, though they were not bound directly by CHOP. The data support a model where CHOP down-regulates hepatocyte metabolism and identity via regulation of ONECUT1. This is a new and interesting finding, perhaps explaining the steatosis phenotype of livers that accompanies ER stress, although this was not tested directly.

      The second main finding of this paper is that CHOP deletion leads to an interesting assortment of effects on genes related to the ER stress response and integrated stress response (ISR). As expected, based on prior work, CHOP deletion led to more phosphorylation of eIF2alpha (CHOP is known to upregulate the phosphatase for this translation factor). However, unexpectedly, this did not cause increased expression of ATF4 (a transcription factor whose upregulation during stress is dependent on eIF2alpha phosphorylation) and its downstream targets; in fact, CHOP deletion had the opposite effect on these. In other words, CHOP seems to both turn off the initiating signal for the ISR (namely, eIF2alpha phosphorylation) and also promote the downstream signaling events that rely on this initiating signal. It makes sense that cells would do this, as restoring translation would be important for realizing the effects of the massive changes in gene expression initiated by ER stress, and yet this would exacerbate stress in the short term, so it would be counterproductive to also turn off the entire stress-regulated program. Having a factor (perhaps CHOP) that coordinates these two events makes sense. It will be interesting in future work to understand the mechanisms behind this regulation.

      Finally, CHOP deletion led to less activity of other aspects of the ER stress response, notably IRE1 (determined through measurement of XBP1 splicing and RIDD of Bloc1s1). This is explained by the continued phosphorylation of eIF2alpha in these knockouts, as the continued attenuation of translation would lessen the burden of misfolded proteins in the ER. Somewhat confusingly, the same pattern is not seen in downstream targets of XBP1. Less splicing, coupled with perhaps less translation of the spliced mRNA, should result in less active transcription factor and lower expression of its target genes in the CHOP KO. This is not observed in Figure 2, although the more global gene expression analysis suggests that all stress-dependent gene expression changes were weaker in the CHOP KO livers.

      The authors characterize the effects of CHOP, promoting restoration of protein synthesis and the accompanying exacerbation of stress while preserving the signaling that should relieve ER stress, as a switch from an acute to chronic phase of ER stress. This is mirrored in their analysis of ATF6 in a similar series of experiments. Although this is an interesting framework for thinking about the stress response, whether CHOP is the key factor or a supporting actor in regulating this transition will require a better understanding of the mechanisms involved.

    5. Author response:

      We thank the reviewers for their assessment of our work and their comments. We are grateful for their evaluation of our findings as fundamental and convincingly supported, and for their appreciation of the relative scope of this manuscript and of future work. The most direct requests for new experimental data are from reviewer #2, who asks for direct assessment of the effects of CHOP deletion on expression of GADD34 and on protein synthesis. We agree that these are important experiments to conduct for the revision.

      The reviewers requested more clarity on the experimental logic of the paper and on the place of our findings in the broader context of ER stress signaling, which we will be happy to provide in a revised manuscript. These revisions will include a more explicit consideration of how the regulation of metabolic genes by CHOP contributes to its effects in the liver independently of its role in regulating eIF2a dephosphorylation.

      In particular, there were concerns about the logic of the time points chosen that we feel are important to also address here. For analysis of ChopHKO animals, all experiments were carried out 8 hours after ER stress challenge. This is because, as we show in Fig. 1B and also in our previous paper on CHOP (1), this is the time point at which CHOP expression is at its maximum. Thereafter, hepatocytes become heterogeneous with respect to whether they do or do not express CHOP. This is an interesting finding because it suggests that CHOP is part of a cellular switch, and potentially even an effector of that switch—a point currently raised in the Discussion but worth further highlighting in a revision. At the practical level, it means that discerning the contribution of CHOP to ER stress signaling and adaptation at subsequent time points will require sophisticated single cell analyses that can discriminate cells that express CHOP from cells that do not, which are an important future direction.

      In contrast, for Atf6aHKO animals, all experiments were carried out 48 hours after ER stress challenge. As we have previously shown (2), at short time points after a stress challenge, such as 8 hours, there is very little difference in ER stress signaling between wild-type animals and those lacking ATF6a. The reason for this lack of distinction is that the major targets of ATF6a are ER chaperones and the like. Because adaptation to ER stress in the early phases of the response depends more on non-transcriptional mechanisms such as inhibition of protein synthesis and IRE1-dependent mRNA decay (RIDD), the failure to fully upregulate ATF6a targets is initially of little consequence. It is only at later time points when wild-type animals restore ER homeostasis and largely silence ER stress signaling. In contrast, at these same later points, animals lacking ATF6a show evidence of persistent ER stress, most notably in the form of persistent Xbp1 mRNA splicing and profound suppression of metabolic genes. Although the 8 hour time point for experiments in ChopHKO animals differs from the 48 hour time point for Atf6aHKO animals, the two lines of experimentation are united by the persistence of ongoing ER stress and of ISR signaling despite diminished eIF2a phosphorylation at the points when the presence of CHOP or the absence of ATF6a are of the greatest impact. A revised manuscript will present this logic more clearly.

      References

      (1)  Liu K, et al., EMBO Reports 25, 228 (2024)

      (2)  Rutkowski DT, et al., Dev. Cell 15, 829 (2008)

    1. eLife Assessment

      This important study combines behavioural analysis, voltage imaging and electrophysiology to advance our understanding of muscle coordination at the cell-to-cell level, in Caenorhabditis elegans. The evidence supporting the conclusions is convincing; however, the use of correlation in some aspects of data interpretation is a relative weakness. The technically sophisticated optogenetic voltage clamp approach introduced here can be applied to other small, transparent animals, making these findings of broad interest to researchers studying electrical coupling between cells or utilising optical electrophysiology techniques.

    2. Reviewer #1 (Public review):

      Summary:

      This study aims to reveal the contribution of individual gap junction proteins to the signal transmission and connectivity of living C. elegans animals in a completely non-invasive way through all-optical electrophysiology. The authors achieve this by simultaneous expression of bipoles, an excitatory/inhibitory light-activated actuator and Quasar2, a genetically encoded voltage dye. With this study, the authors extend their previous efforts to leverage the strength of optogenetic neurophysiology and set a new standard in this domain. In addition, they adapted their established methods to perform cell-specific optogenetic voltage clamp and revealed changes in gap junction connectivity. They also find that increasing excitability in innexin mutants is indicative of a reduction in gap-junction connectivity and current leaks.

      Strengths:

      This is an extremely strong manuscript, a technical feat and tour de force to infer junctional coupling through all-optical electrophysiology. The establishment of the voltage clamp method is powerful and allows researchers to obtain not only tight control over voltage signals but also permits the investigation of gap junction function in response to positive and negative voltage steps in a completely non-invasive fashion. This will be a new paradigm for investigating muscle electrophysiology in future.

      Weaknesses:

      This is a strong pioneering study, and I found very few technical weaknesses. The correlation quantification is relatively weak to establish connective causality, as a shared upstream input may lead to a similar perceived correlation. This is especially concerning for an average lag time of ~0, and the authors may want to investigate if there is unchanged connectivity in an unc-31 or unc-13 mutant. Conceptually, the local connectivity is scaled to account for behaviour: future studies may wish to perform this method on moving animals, and in specific neuronal populations, where a non-invasive optogenetic voltage clamp method will truly shine.

    3. Reviewer #2 (Public review):

      Summary:

      This technically sophisticated study combines behavioral analysis, voltage imaging, electrophysiology, and a newly developed cell-specific optogenetic voltage clamp (cOVC) approach to investigate gap-junction (GJ)-mediated coupling in C. elegans body-wall muscle cells. The work explores the coordination of muscle cells and systems physiology and introduces a method with potential utility beyond the nematode system studied.

      Strengths:

      The main strength of the work is the development and application of the cOVC method. This approach enables minimally invasive in vivo assessment of cell-to-cell electrical coupling in intact animals. This technique represents a meaningful advance over traditional electrophysiological techniques that require dissection or cell isolation.

      With respect to the GJ biology and function, the authors support their conclusions by integrating additional independent experimental approaches. Findings from behavioural/locomotion assays, voltage imaging, patch-clamp recordings, and cOVC measurements are generally consistent, particularly for unc-9 mutants, which show reduced synchronization of muscle cells, impaired electrical coupling, and severe locomotor defects.

      The gain-of-function experiment using murine Cx36 suggests that more or less electrical coupling can disrupt (normal) locomotion.

      Weaknesses:

      The main issue of this otherwise excellent manuscript relates to interpretation rather than experimental quality. Throughout the manuscript, increased correlation is often interpreted as evidence of increased electrical coupling. Are correlation, synchrony, and conductance equivalent measures? If not, how would this affect these correlations? Furthermore, could broader action potentials and altered excitability also increase correlation values? This concern could be addressed through a discussion of this limitation.

      Similarly, the proposed mechanism that reduced GJ coupling increases excitability through reduced leak currents is plausible but not directly demonstrated. Are alternative explanations, e.g., compensatory changes in ion-channel expression or gap-junction composition, possible? These could also be considered to improve the balance of this work.

      The conclusions regarding Cx36 overexpression would also benefit from more cautious wording, as developmental or localization effects have not been excluded.

      However, the experimental dataset is very strong. In my opinion, no major additional studies are needed. Direct analysis of compensatory changes in innexin expression or localization could strengthen the interpretation of the proposed mechanism. Overall, the study is of high technical quality, contains a notable methodological advance, and provides important insights into muscle synchronization and GJ biology.

    1. eLife Assessment

      This important study describes how selective lesions of key cortical and subcortical motor areas affect reaching actions in macaque monkeys. The results will be of interest to both basic and clinical researchers studying the neural control of movement. Kinematic analysis of movement quality is solid but could be improved by considering other metrics, especially those that relate to grasping. Evidence for the general claims related to the role of specific motor areas is incomplete because the lesions did not fully eliminate any single area while simultaneously involving neighbouring areas.

    2. Reviewer #1 (Public review):

      Summary:

      This is a very interesting and well-done study of the effects of selective lesions to the sensorimotor cortex and the red nucleus on control of upper limb movements. The findings that the red nucleus may subserve recovery of upper limb motor function after cortical lesions in macaques and the different motor functions of different cortical sensorimotor areas are significant findings of considerable interest to sensorimotor neuroscientists, neurologists and neurosurgeons. The methods are mostly excellent, but there are some questions about the use of endothelin lesions in cortical areas and the use of trajectory variability as a marker of movement quality and fine motor control. Furthermore, it is questionable that increased trajectory variability in reaching a target reflects reduced movement quality, reduced ability to independently control muscles, and is a proximal analog of reduced dexterity.

      Strengths:

      The rationale that rubrospinal projections onto spinal neurons may subserve the good recovery of upper limb movements observed after lesions of sensorimotor cortex is compelling. The methods involving complete lesions of the red nucleus followed by recovery prior to lesions affecting various sensorimotor cortical areas are a strength. The excellent interpretations offered in the Discussion section are also a strength.

      Weaknesses:

      There are weaknesses in the Methods, including:

      (1) no information on dimensions of the cup containing the food reward or types of food rewards,

      (2) recording 3D hand movements with a single camera,

      (3) cortical endothelin lesions were not very precise,

      (4) the use of trajectory variability as a measure of movement quality and reduced ability to independently control muscles.

      Some interpretations presented in the Discussion are not well supported. The discussion related to movement quality should be modified to focus on trajectory variability. The suggestion that rubrospinal projections onto motor neurons are apparently irreplaceable is not well justified because one monkey receiving a complete red nucleus lesion showed nearly full recovery of maximum movement speed, while the other monkey did not. The nearly full recovery of one monkey was probably due to new corticospinal connections onto motor neurons, whereas it is possible that the other monkey would have recovered better given more time before the 2nd lesion to cortical areas.

    3. Reviewer #2 (Public review):

      Summary:

      This study made selective lesions in motor cortical subregions and the magnocellular red nucleus in nine macaque monkeys, and evaluated reaching and grasping movements using maximum speed and trajectory variability. The results suggest that damage to the posterior old primary motor cortex (M1) was mainly associated with reduced maximum speed, whereas damage to the new M1 was mainly associated with increased trajectory variability. Damage to the anterior old M1 did not clearly add further impairment. Lesions of the magnocellular red nucleus (RNm) alone mainly reduced reaching speed, but recovery after subsequent cortical lesions was worse than after cortical lesions alone, suggesting that the rubrospinal pathway may be important for compensation after cortical damage in monkeys. Overall, this is a valuable study that examines differences among M1 subregions using selective lesions in macaques.

      Strengths:

      (1) This study tackles an important question. It attempts to decompose the diverse upper-limb impairments after stroke into the effects of different primary motor cortex subregions.

      (2) Another strength is that the lesions and behavioral impairments were evaluated quantitatively. The use of nine macaque monkeys with different lesion patterns, together with quantitative behavioral evaluation, provides a rare and valuable dataset. The authors also followed recovery using quantitative behavioral measures such as maximum speed and trajectory variability.

      (3) The inclusion of RNm lesions is also valuable, as it revisits the classic question raised by Lawrence and Kuypers (1968) of how brainstem descending pathways contribute to recovery after cortical motor damage.

      Weaknesses:

      (1) The main limitation is that the contribution of each cortical lesion is sometimes interpreted from largely qualitative comparisons. Because the lesion extent was not always limited to the intended region, it is difficult to fully separate the independent contribution of each subregion. Some conclusions are also based on comparisons between a small number of animals. The dataset itself is valuable, and the manuscript would be strengthened by presenting these conclusions more cautiously and explicitly acknowledging this limitation.

      (2) Because the behavioral evaluation is quantitative, it would be helpful to show the relationship between lesion size and behavioral impairment more quantitatively. For example, rank correlations between the lesion size of each cortical region and behavioral measures could help readers evaluate whether the type and size of lesion are related to behavioral impairment.

      (3) The discussion of area 4s could be further developed. The authors suggest that this region may have a different role, but the specific hypothesis is not fully clear. There has also been skepticism in the previous literature about area 4s, for example, Meyers et al. (1954), and this broader background could be discussed. (Meyers R, Knott JR, Skultety FM, Imler R (1954) On the Question as to the Existence of a "4s" Suppressor Mechanism. Journal of Neurosurgery 11:7-23.)

    4. Reviewer #3 (Public review):

      Summary:

      In this article, the authors performed targeted lesions in cortical areas involved in forelimb control of rhesus macaques. Using a reaching task with kinematic tracking, they compared kinematic variability (as a proxy for dexterity) and reaching speed (as a proxy for strength) before and after cortical (n=7) and magnocellular red nucleus (RNm) (n=2) lesion. Changes in these movement metrics were related to the location and extent of the lesions, reconstructed from histology. The authors report that lesions with a large component in New M1 had a pronounced effect on kinematic variability, whereas lesions with a large component in posterior Old M1 primarily affected reaching speed. Lesions of the RNm were performed in two animals approximately seven to eight weeks before the cortical lesion. By themselves, RNm lesions produced a significant but small reduction in reach speed. They also magnified the effect of the subsequent cortical lesions.

      Strengths:

      (1) For non-human primate (NHP) research, this is a large cohort.

      (2) The behavioural analyses are clear and precise.

      (3) The additional red nucleus lesions in two monkeys provide unique complementary information.

      Weaknesses:

      (1) Description of injuries. As described and reported in the current result section and figures, readers do not have a clear understanding of the lesion extent and location.

      (2) Lack of formal correlative analyses between lesion characteristics and behaviour. Currently, it seems that the conclusions are based on impressions between some aspects of the lesions and precise kinematic measures.

      (3) The data will be of interest to a large community of researchers working on brain injury and stroke recovery, as well as cortical motor control. There are, however, some major methodological issues in the current version of the manuscript that prevent a clear evaluation of the findings and potential contribution of the work in the field. These issues need to be addressed to support the conclusions.

    1. eLife Assessment

      The framework of this potentially important study – with the integration of multiple levels of analysis, glymphatic MRI, transcriptomics, functional MRI, and public amyloid maps, in one framework – is clever. The assertion that regional amyloid vulnerability may depend not just on neural activity alone, but on whether clearance is appropriately matched to activity, is an interesting and novel concept. However, the chosen approach to imaging glymphatic clearance relies on indirect inferences from a small subgroup. In its current form, the main conclusions of this study are therefore incompletely supported.

    2. Reviewer #1 (Public review):

      Summary:

      Regional differences in the brain's waste-clearance system may interact with neural activity to influence where amyloid-B accumulates. Using intrathecal GBCA administration to produce "Glymphatic MRI" in 96 subjects, the authors mapped cortical glymphatic influx and clearance and found distinct spatial patterns, with transcriptomic analyses linking better glymphatic function to neuronal cell types (through genes). In a subgroup with resting-state fMRI, regions with stronger resting-state activation generally showed higher contrast clearance, indicating a positive coupling between these processes. Notably, cortical regions where neural activity and glymphatic clearance were mismatched showed greater amyloid-β burden in a separate, publicly available PiB-PET dataset, suggesting that activity-clearance decoupling may contribute to regional vulnerability and neurodegeneration.

      Strengths:

      This is a rare and valuable dataset. Intrathecal contrast injection in ~100 subjects is quite a remarkable accomplishment alone, but the addition of resting-state fMRI, a correlative PiB cohort, and gene-expression pattern data is impressive.

      Weaknesses:

      This is a cross-sectional study, and we can't determine whether neural activity drives glymphatic clearance, whether glymphatic dysfunction alters neural activity, or whether both are shaped by a third factor. Language describing "flow", "influx", and "clearance" could be made more specific so the reader can more easily follow the methodological approach.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Li et al. investigated the relationships among regional cortical tracer dynamics following intrathecal gadolinium administration, neural activity, and amyloid-β deposition in humans. Using serial MRI acquisitions after intrathecal gadodiamide administration in 96 participants, the authors characterized regional signal enhancement and clearance patterns across the human cortex. They integrated these imaging measures with transcriptomic data (Allen Human Brain Atlas), resting-state fMRI outcomes, and an external amyloid PET dataset. The authors report that regions with more efficient tracer clearance are enriched for genes related to synaptic organization and neuronal cell types, that tracer clearance patterns are in parts spatially coupled to spontaneous neural activity, and that regional mismatch between neural activity and tracer clearance is associated with increased amyloid burden according to the PET dataset.

      Strengths:

      The study addresses an important and very timely question about the interaction among neural activity, cerebrospinal fluid dynamics (waste clearance), and regional vulnerability to neurodegeneration. Integrating serial post-contrast MRI, transcriptomics, resting-state fMRI, and amyloid imaging is ambitious and conceptually very interesting. The spatial characterization of cortical tracer dynamics is potentially valuable for the field, particularly given the increasing interest in human glymphatic imaging approaches and intrathecal contrast MRI, which provides an opportunity to assess CSF tracer dynamics without confounding tracer signal from the blood. The imaging preprocessing pipeline includes normalization of regional cortical signal intensity to a reference region within each session before calculation of longitudinal percentage change, which helps reduce inter-session variability within individuals for conventional T1-weighted imaging. The transcriptomic analyses linking tracer dynamics to neuronal and synaptic gene expression patterns are also interesting. In addition, the manuscript addresses recent literature on neurovascular coupling, glymphatic function, and amyloid vulnerability.

      Weaknesses:

      Several issues limit the strength of the conclusions. One concern relates to the interpretation of repeated post-intrathecal contrast MRI measurements as direct indicators of glymphatic influx and clearance. The approach presented by the authors measures regional signal changes following intrathecal gadodiamide administration, but does not directly visualize paravascular flow or establish that the observed signal dynamics specifically reflect glymphatic transport mechanisms. Although it is widely accepted that CSF influx occurs primarily along periarterial spaces as part of the glymphatic system, and the terminology "glymphatic MRI" is increasingly used in the literature, the physiological processes contributing to delayed parenchymal enhancement, including CSF-interstitial exchange mediated by convective bulk flow and/or extracellular diffusion, as well as transient and, in the case of linear gadolinium agents, even long-term tracer retention remain incompletely resolved. Importantly, tracer kinetics may not directly reflect interstitial fluid kinetics, as solute transport may also be influenced by compartmental and extracellular barriers, diffusion constraints, and tissue retention effects. As currently written, several sections of the manuscript appear to overstate what can be directly inferred from the imaging data. This issue may be particularly relevant given the intrathecal use of gadodiamide (Omniscan), a linear gadolinium-based contrast agent with known long-lasting tissue retention due to lower kinetic stability compared to macrocyclic agents. Sustained signal at later imaging time points may therefore not only reflect impaired glymphatic clearance dynamics may also be influenced by tissue retention of contrast material, particularly in the context of neurological disease. In addition, the participant cohort is heterogeneous and includes individuals with neuroinflammatory and neurodegenerative diseases, peripheral neuropathy, and motor neuron disease. Although the authors argue that the spatial tracer patterns are relatively preserved across neurodegenerative groups, this heterogeneity complicates interpretation of imaging data and raises the possibility that disease-related factors and altered tracer-tissue interactions contribute to the observed effects. Thus, the rationale for interpreting a greater tracer signal at 39h as evidence of impaired glymphatic clearance should be explained more carefully, particularly given the highly heterogeneous patient population.

      In addition, the analyses linking spontaneous neural activity and tracer clearance are based on a very small rs-fMRI subgroup (n = 15), limiting the generalizability. The interpretation of the "mismatch" analysis also requires caution. The mismatch index was computed from z-scored fALFF and tracer clearance and is subsequently associated with amyloid burden derived from the external PET dataset rather than from the studied participants themselves. Therefore, the observed spatial associations should be interpreted with greater caution rather than as evidence for a direct mechanistic relationship. The cross-sectional nature of the analyses also limits conclusions regarding the directionality and temporal sequence of the relationships between neural activity, tracer dynamics, and amyloid burden. Several statements in the Discussion currently imply stronger causal or biological conclusions than are directly supported by the data.

      Despite these limitations, the study presents an interesting dataset and proposes a framework for understanding regional vulnerability to protein accumulation in neurodegeneration. This work hopefully motivates further investigation into the important relationships among neural activity, CSF dynamics, and neurodegeneration in humans.

    4. Reviewer #3 (Public review):

      This manuscript addresses an interesting and timely question: whether regional glymphatic clearance in the human cortex is spatially coupled to neural activity and whether a mismatch between activity and clearance may help explain regional vulnerability to amyloid-β deposition. The authors use intrathecal gadolinium-based glymphatic MRI in 96 participants, derive cortical influx and clearance maps, integrate these with Allen Human Brain Atlas transcriptomic data, and then relate regional clearance to resting-state fMRI measures in a smaller subgroup. They further compare the resulting activity-clearance mismatch map with an open-source ¹¹C-PiB amyloid PET dataset. The overall concept is attractive because it attempts to connect glymphatic physiology, neuronal activity, and proteopathy at the regional level of the human brain, an important and understudied area.

      The main strength of the study is the use of direct intrathecal contrast-enhanced MRI to generate cortical maps of glymphatic tracer dynamics. This is a technically demanding approach and provides a richer spatial readout than indirect MRI proxies of glymphatic function. The authors show that the cortical tracer signal increases from 4.5 h to 15 h and then decreases by 39 h, allowing them to interpret the early signal as reflecting influx and the persistent signal at 39 h as impaired clearance. They further identify regional patterns, with faster influx in medial prefrontal/insular areas and slower clearance in dorsal prefrontal and parietal surface regions. The analysis is visually clear, and the use of cortical gradients is a useful way to reduce complex regional data into interpretable spatial axes.

      The multimodal integration is also interesting. The transcriptomic analysis suggests that regions with faster glymphatic clearance are enriched for synaptic organisation and neuronal activity-related pathways, while regions with slower clearance show enrichment for metabolic and mitochondrial pathways. The cell-type enrichment analysis further implicates excitatory and inhibitory neurons, oligodendrocyte lineage cells, microglia and, to a lesser extent, astrocytes. This provides a plausible biological bridge between regional neural activity and clearance function, and the sensitivity analysis using ReHo in addition to fALFF is a useful robustness check.

      However, the manuscript should be more careful in its causal interpretation. The study is cross-sectional and largely correlative in space. The finding that regions with higher spontaneous neural activity tend to show better glymphatic clearance is intriguing, but it does not establish that neural activity drives clearance in these participants. Conversely, it remains possible that better tissue integrity, vascular function, CSF access, cortical geometry, vascular density, or disease composition jointly influence both fMRI measures and tracer clearance. The authors do acknowledge some of these limitations, but the abstract and discussion should more consistently frame the findings as associations rather than evidence of an activity-clearance mechanism in humans.

      The most important limitation is the small size of the fMRI subgroup. Although the whole glymphatic MRI cohort includes 96 participants, the key activity-clearance analysis is based on only 15 individuals, including 11 with peripheral neuropathy and 4 with motor neuron disease. This is a very small and clinically heterogeneous sample on which to build a central conclusion about regional neural activity and glymphatic clearance. The authors show that the 39 h PC map in the fMRI subgroup resembles the whole-cohort map, which is helpful, but this does not address whether the fALFF-clearance relationship is robust at the individual level. The paper would be strengthened by reporting subject-level stability, leave-one-out analyses, and whether the association persists after excluding the four motor neuron disease cases.

      A second major concern is the interpretation of the amyloid analysis. The ¹¹C-PiB map is derived from an external open-source Alzheimer's disease dataset, not from the same participants who underwent glymphatic MRI and fMRI. Therefore, the association between activity-clearance mismatch and amyloid burden is a spatial correspondence across group-average maps, not an individual-level relationship. This is valuable for hypothesis generation, but should not be presented as evidence that a mismatch in the present cohort predicts amyloid deposition. The authors should clearly state that this analysis tests whether mismatch regions overlap with known amyloid-prone cortical regions, rather than directly linking mismatch to amyloidosis in individual participants.

      The definition of "mismatch" also needs clarification. The text defines the mismatch index as the negative absolute difference between z-fALFF and z-39h PC, and states that higher scores indicate greater mismatch. Because the index is negative, values closer to zero would normally indicate a smaller absolute difference rather than a greater mismatch. This should be checked carefully and corrected if necessary. More broadly, because a higher 39 h PC indicates worse clearance, the interpretation of match and mismatch categories is not intuitive. The authors should provide a clearer schematic and ensure that the mathematical definition, biological interpretation and figure labelling are fully aligned.

      Several technical confounds require more attention. Intrathecal gadolinium MRI is influenced by CSF dynamics, posture, sleep, circadian timing, renal clearance, age, intracranial pathology, and potentially diagnosis-specific differences. The authors acquired scans at fixed time points and noted that patients slept as usual, but individual sleep duration, sleep quality, posture, and daytime activity were not objectively measured. Given that the central claim concerns glymphatic clearance, these are not minor confounders. The authors should consider adjusting for age, sex, diagnosis, vascular risk factors, and relevant clinical variables where possible, and be more explicit about how heterogeneous disease indications may influence cortical tracer kinetics.

      The statistics are generally good. However, many correlations are performed across 400 cortical parcels, which are not independent biological samples. The paper would benefit from clearer separation between participant-level inference and region-level spatial inference. For example, the fALFF-clearance and mismatch-amyloid analyses are regional map correlations, not correlations across individuals. This should be clearly stated throughout. The authors should also report effect sizes and confidence intervals more consistently, and explain how multiple comparisons were controlled across transcriptomic, cell-type, fMRI, ReHo and amyloid analyses.

      The transcriptomic analysis is useful but should be presented as indirect. AHBA data come from six post-mortem brains; only the left hemisphere was used, and the donors were healthy and younger than the clinical cohort. Therefore, these data capture intrinsic regional gene-expression patterns rather than disease-state expression in the same individuals. The authors should avoid implying that the transcriptomic findings directly explain glymphatic function in their participants. The current discussion partly acknowledges this, but the framing in the abstract and results could be more cautious.

      There are also several points of presentation that should be improved. The manuscript should consistently distinguish glymphatic influx, glymphatic clearance, CSF tracer retention, and waste clearance. A 39 h residual gadolinium signal is a useful proxy for delayed clearance, but it is not the same as direct measurement of amyloid or tau clearance. The language around "waste clearance" and "amyloidosis" should therefore be precise. The authors should also clarity whether "higher clearance" corresponds to lower 39 h PC across all analyses, as this inversion is easy for readers to misinterpret.

    5. Author response:

      We would like to express our deepest gratitude to the Editors and Reviewers for their highly rigorous and constructive evaluation of our manuscript. We are greatly encouraged by the recognition of our study’s ambition, the unique value of the in vivo intrathecal contrast MRI dataset, and the conceptual novelty of linking macroscopic glymphatic physiology with neural activity and regional proteopathy.

      We fully agree with the thoughtful limitations and methodological concerns raised in the eLife Assessment and the Public Reviews. In our upcoming revised manuscript, we are implementing a comprehensive set of revisions to address these points. Specifically, our planned revisions focus on the following key areas:

      - Tempering Causal Interpretations: We agree that our cross-sectional design precludes definitive causal inferences. We are systematically revising the manuscript to soften causal language (e.g., replacing "drives" with "is spatially associated with"). We will explicitly frame our findings as macroscopic spatial associations and discuss the potential influence of joint physiological confounders.

      - Tightening Terminology and Imaging Physics: We are refining our terminology to more accurately reflect our MRI measurements. We will replace assertive terms like "direct glymphatic flow" with precise descriptors such as "imaging proxies for tracer enhancement and retention." Furthermore, we are expanding the Limitations section to explicitly acknowledge the confounding effects of Partial Volume Averaging (PVE), systemic tracer redistribution, and renal clearance kinetics.

      - Conducting Supplementary Imaging & Robustness Analyses: To address concerns regarding cohort heterogeneity and the sample size of the rs-fMRI subgroup (n=15), we are performing a series of rigorous supplementary analyses. This includes conducting sensitivity analyses (e.g., excluding the motor neuron disease subgroup) and applying leave-one-out cross-validation to rigorously assess the subject-level stability and robustness of the spatial coupling between neural activity and tracer clearance.

      - Clarifying the Conceptual Model and "Mismatch" Index: To improve readability, we are moving the anatomical definitions of the cortical gradients directly into the Results section. Additionally, we are introducing schematic diagram to intuitively explain the mathematical formulation and biological interpretation of the "activity-clearance mismatch" index.

      - Re-framing External Dataset Analyses: We are carefully re-framing the interpretations of the Allen Human Brain Atlas (AHBA) transcriptomic data and the external PiB-PET amyloid dataset, emphasizing that these reflect spatial correspondences of intrinsic regional vulnerability across groups, rather than individual-level direct interactions.

      We believe these revisions will significantly enhance the scientific rigor, clarity, and precision of our study.

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      We thank the reviewer for the kind comments.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      We have added more details on the annotation process and have included the list of instructions sent to the annotators. We have also included mmpose configs with the code provided. The following files include the relevant details:

      File including the list of instructions sent to the annotators:

      OpenMonkeyWild Photograph Rubric.pdf

      Mmpose configs:

      i) TopDownOAPDataset.py

      ii) animal_oap_dataset.py

      iii) init.py

      iv) hrnet_w48_oap_256x192_full.py

      Anaconda environment files:

      i) OpenApePose.yml

      ii) requirements.txt

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths:

      Open source dataset with excellent annotations on the format, as well as example code provided for working with it.

      Properties of the dataset are mostly well described.

      Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones?

      Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out).

      The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      We thank the reviewer for the kind comments.

      Weaknesses:

      More details on training hyperparameters used (preferably full config if trained via mmpose).

      We have now included mmpose configs and anaconda environment files that allow researchers to use the dataset with specific versions of mmpose and other packages we trained our models with. The list of files is provided above.

      Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010).

      We have included a datasheet for our dataset in the appendix lines 621-764.

      Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here.

      We have included the list of instructions sent to the Hive annotators in the supplementary materials. File: OpenMonkeyWild Photograph Rubric.pdf

      Should include model cards, as described in Mitchell et al (arXiv:1810.03993).

      We have included a model card for the included model in the results section line 359. See Author response image 1:

      Author response image 1.

      It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year.

      We agree that the source could introduce structural biases. This is why we included images from so many different sources and captured images at different times from the same source—in hopes that a large variety of background and lighting conditions are represented. However, doing so limits our ability to document each source background and lighting condition separately.

      Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically.

      The focus of this work is on overall keypoint localization accuracy and hence we wanted a metric that is easy to interpret and implement, in this case we made use of PCK (Percentage of Correct Keypoints). PCK is a simple and widely used metric that measures the percentage of correctly localized keypoints within a certain distance threshold from their corresponding groundtruth keypoints.

      A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

      We have now included a histogram of unnormalized bounding boxes in the manuscript, see Author response image 2:

      Author response image 2.

      Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential.

      We thank the reviewer for the kind comments.

      However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      We thank the reviewer for pointing this out. Inter-observer reliability is important for ensuring the quality of the annotations. We first used Amazon MTurk to crowd source annotations and found that the inter-observer reliability and the annotation quality was poor. This was the reason for choosing a commercial service such as Hive AI. As the crowd sourcing and quality control are managed by Hive through their internal procedures, we do not have access to data that can allow us to assess inter-observer reliability. However, the annotation quality was assessed by first author ND through manual inspections of the annotations visualized on all of the images the database. Additionally, our ablation experiments with high out of sample performances further vaildate the quality of the annotations.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      Our goal was to obtain as many images as possible from the most commonly studied ape species. In order to ensure a large enough database, we focused only on the species and combined images from as many sources as possible to reach our goal of ~10,000 images per species. With the wide range of people involved in obtaining the images, we could not ensure that all the photographers had the necessary expertise to differentiate individuals and subspecies of the subjects they were photographing. We could only ensure that the right species was being photographed. Hence, we cannot include more detailed information.

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      We created the training set to reflect the species composition of the whole dataset, but used test sets balanced by species. This was done to give a sense of the performance of a model that could be trained with the entire dataset, that does not have the species fully balanced. We believe that researchers interested in training models using this dataset for behavior tracking applications would use the entire dataset to fully leverage the variation in the dataset. However, for those interested in training models with balanced species, we provide an annotation file with all the images included, which would allow researchers to create their own training and test sets that meet their specific needs. We have added this justification in the manuscript to guide the other users with different needs. Lines 530-534: “We did not balance our training set for the species as we wanted to utilize the full variation in the dataset and assess models trained with the proportion of species as reflected in the dataset. We provide annotations including the entire dataset to allow others to make create their own training/validation/test sets that suit their needs.”

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

      We thank the reviewer for this important suggestion. Since we can separate a video into its constituent frames, one can indeed use the provided model or other models trained using this dataset for inference on the frames, thus allowing video tracking applications. We now include a short video clip of a chimpanzee with inferences from the provided model visualized in the supplementary materials.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Please provide a more thorough description of the annotation procedure (i.e., the instructions given to crowd workers)! See public review for reference on dataset annotation reporting cards.

      We have included the list of instructions for Hive annotators in the supplementary materials.

      An estimate of the crowd worker accuracy and variability would be super valuable!

      While we agree that this is useful, we do not have access to Hive internal data on crowd worker IDs that could allow us to estimate these metrics. Furthermore, we assessed each image manually to ensure good annotation quality.

      In the methods section it is reported that images were discarded because they were either too blurry, small, or highly occluded. Further quantification could be provided. How many images were discarded per species?

      It’s not really clear to us why this is interesting or important. We used a large number of photographers and annotators, some of whom gave a high ratio of great images; some of whom gave a poor ratio. But it’s not clear what those ratios tell us.

      Placing the numerical values at the end of the bars would make the graphs more readable in Figures 4 and 5.

      We thank the reviewer for this suggestion. While we agree that this can help, we do not have space to include the number in a font size that would be readable. Smaller font sizes that are likely to fit may not be readable for all readers. We have included the numerical values in the main text in the results section for those interested and hope that the figures provide a qualitative sense of the results to the readers.

    2. Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential. However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

    3. Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths: - Open source dataset with excellent annotations on the format, as well as example code provided for working with it. - Properties of the dataset are mostly well described. - Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones? - Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out). - The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      Weaknesses: - More details on training hyperparameters used (preferably full config if trained via mmpose). - Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010). - Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here. - Should include model cards, as described in Mitchell et al (arXiv:1810.03993). - It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year. - Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically. - A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

    4. eLife Assessment

      The OpenApePose dataset presented in this manuscript represents an important contribution to the field of primate behaviour and computer-vision science with methodological applications that are sure to be applicable for a wide variety of taxa. The analysis supporting the utility of this database is solid and compelling but would benefit from some additional clarity, particularly with regards to the annotation of landmarks, model parameters and division of the dataset for training, validation and testing.

    1. eLife Assessment

      This study addresses an important question in aging biology by combining metabolomics, transcriptomics, molecular genetics, and functional analyses to examine how cytosolic acetyl-CoA metabolism influences late-life fitness in replicatively aging yeast. The evidence supporting the roles of AMPK activation, mitochondrial acetyl-CoA utilization, and fatty acid synthesis in preserving fitness during aging is convincing overall, and the engineered A2A strain provides an elegant demonstration that coordinated modulation of distinct acetyl-CoA metabolic branches can increase the proportion of aged cells with a low-senescence phenotype. The study provides significant insight into mechanisms that allow aging cells to maintain fitness without extending replicative lifespan.

    2. Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fuelled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. The reviewers have achieved their aims and the results support their conclusions. This work provides important insight into molecular mechanisms that allow aging without loss of fitness and will be of interest to scientists in the field of aging and metabolism.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells. Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metabolic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Comments on revised version.

      I am fine with the revisions.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fueled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. This work will be of interest to scientists in the field of aging and metabolism. Some clarifications in the text would address the following concerns, which would increase the impact of the study:

      (1) What does activation of AMPK (via PGDP-Sak1 expression) do to the replicative lifespan? How many bud scars, in general, do the subpopulations that are older - yet have less Tom70 (increased mitochondrial fitness) - have, after the 48 hrs timepoint that they are examining? How many divisions occurred in this 48hr time period - i.e. is it long enough to have all cells reach the end of their replicative lifespan? This information is important to rule out that a subset of the mutant cells just divided faster and hence had more divisions within 48 hrs (growing faster and living longer are different things). Having identical growth curves doesn't indicate per se that they all divide at the same rate, as there may be a subpopulation that divides faster and a subpopulation that doesn't grow so well.

      Increasing AMPK activity increases replicative lifespan [PMID: 25869125], but given our finding that AMPK activation splits the population, such replicative lifespan assays are hard to interpret. Bud scar counts have a similar issue. Hence we restricted the lifespan and bud scar analyses to wt and A2A which are more homogenous (Figures S2 B and E). A2A cells at 48 h have ~25% more bud scars than wt cells. Yes, by 48 h most of the cells have lost viability (Figure 2E). The reviewer is correct that you can't properly compare the lifespan curves if the cells divide at different rates, hence our follow-up test of wt at 48 h vs A2A at 40 h viability after we had confirmed that these time points captured cells at equivalent replicative ages (Figure 2D, E). This shows that viability of A2A is slightly lower than wt at matched age, indicating a slightly shorter lifespan. 

      (2) A2A cells do not have an extended replicative lifespan (RLS) but show an increase in the "low senescence" population (Figure 2). If the cells are not becoming senescent, why don't they have longer RLS? Not having a longer lifespan seems inconsistent with the statement that "bud scar counting confirmed that A2A cells reach a higher age than wild type", which comes back to how many times the cells can divide in the 48hr timepoint studied and their rate of cell division? Also, the lifespan curve shown is plotted against time, not cell division number, which does not take into account different division times of cells within the population (described above). It would be much more useful to show standard lifespan curves showing cell division numbers per lifespan per cell.

      Our observation that cells can reach the end of life without senescing is consistent with other studies that have studied the life course of individual cells by microscopy [PMID: 31291577, 32675375]. These studies always highlight some proportion of the cells that reach the end of life with no or minimal senescence, though this fraction varies with the experimental system. The question of why cells lose viability without senescing is a complete unknown in the field, and reflects a wider lack of consensus as to why yeast lose viability with replicative age.

      In liquid culture we can only assess viability over time, not cell division number, which we agree is not optimal and we are wary about making strong statements on lifespan for exactly the reasons the reviewer notes. Unfortunately, it is clear from the comparison of liquid and solid media lifespans performed by the Gottschling lab [PMID: 19652178] that culture system has a huge effect on lifespan, with cells in classical plate-based microdissection assays living far longer than the same strains do in liquid. This means that lifespans determined by microdissection-based assays are of questionable relevance to ageing studies performed in liquid culture. Senescence cannot be assayed on plates, while microfluidic systems lack the throughput necessary and preclude key techniques like RNA-seq, so liquid culture assays were the only option for this work. We agree that this leaves an unsatisfactory approximation for lifespan measurements, but we consider it critical that everything is measured in the same system. We therefore restricted our conclusion on lifespan to simply say that lifespan of A2A cells is not extended which our data in Figures 2D, E, S2B does support (see also answer to Q1), and therefore with the majority of A2A cells showing low senescence marks and high fitness at 48 h we can conclude that lifespan and fitness loss must be separable.

      We have added a note of these limitations of lifespan measurements in the materials and methods section of the manuscript.

      (3) Increased "fitness" of the old cells is implied from the increased size of the colonies that the old cells can make. However, this is a measure of the fitness of the daughters per se, not the old mother cells. Are the old mothers just passing on healthier mitochondria and more lipids to the daughters, such that they can divide more times? If the aged cells have an "increased fitness", why don't they divide more times themselves (i.e. live longer?).

      Yes, colony growth speed is defined by daughter cell replication, but as long as the daughters and subsequent generations divide at the same rate irrespective of whether they come from a young or old mothers then the size of the colony after 24 hours varies based on the time it took the initial mother to produce a daughter. This is what the assay really measures. We note that aged wildtype mothers often do not divide at all in the first 24 hours after being put on an agar plate (hence the tiny reported colony size), even though they do eventually produce a daughter which then forms a colony, whereas A2A cells tend to produce the first daughter rapidly whether young or old. It is known that daughters of aged wildtype mothers also divide slower, as to some extent do grand-daughters (PMID: 2644196), which will also contribute to differences in colony size, and this may well result from a lipid and/or mitochondrial contribution, but the primary driver of colony size in 24 hours is the time the mother took to initially divide. We have added this detail to the materials and methods section of the manuscript.

      As noted above, the mechanistic basis of lifespan is unknown, but although senescence can shorten lifespan, our work and that of others shows that lifespan is still limited in the absence of senescence.

      (4) The statement is made that "these experiments define two classes of aging cells with distinct metabolic needs, coherent with the model of two aging trajectories previously proposed (referencing Nan Hao's work)". However, the big difference here is that in Nan Hao's work, their two aging trajectories influenced the length of lifespan, but that does not appear to be the case here. That distinction should be made clear. Perhaps the authors could also speculate as to why the A2A yeast stops dividing after presumably the same number of cell divisions, even though they have an activated AMPK and activated fatty acid synthesis pathway.

      Yes, this is a good point and we have added this distinction to the Discussion:

      “Here we have characterised two classes of ageing cells seemingly differentiated by high and low availability of cytosolic Acetyl-CoA, consistent with a previous demonstration that ageing follows two trajectories in yeast though it should be noted that in this previous report, the two trajectories also differed in replicative lifespan (6).”

      We would love to speculate on why the A2A cells don't have an extended lifespan, but at this point we don't have a strong hypothesis. We have come up with many theories for this, but none that we haven’t managed to disprove experimentally. One thing worth considering is that many cells which lose replicative viability in liquid culture and probably in plate assays remain intact – for example, DNA and RNA integrity is not compromised over 24- 48 h – so those cells are probably not dead per se. But we also detect apoptosis-sized DNA fragments, which must come from dead cells, so there is clearly not a single mechanism defining the end of replicative lifespan.

      (5) I am a bit confused by the use of the word "senescence" by this lab here and in their previous growth on galactose studies. If yeast don't senesce, which is usually defined as an irreversible arrest of the cell cycle where cells stop dividing, shouldn't the yeast that do not senesce still be dividing and hence have a longer lifespan? Should a different term be used rather than senescence? Such as "fitness late in life". The authors giving their definition of senescence may help reduce this apparent contradiction.

      We completely agree, this is confusing and noted this distinction in the Introduction. Use of the term senescence to mean a loss of fitness late in life in yeast stems from the classical definition of senescence as applied to whole organisms. However, the term senescence as applied to cells has a more specific meaning in terms of the cell cycle as the reviewer notes. As an individual S. cerevisiae is both a cell and an organism, the terminology clashes. However, the marker we largely employ (Tom70-GFP) which in our hands is a very good proxy for fitness was originally defined as marking the senescence entry point (SEP), so overall we feel we can't avoid the term.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells.

      Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metab to olic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Weaknesses:

      (1) While the manuscript focuses on mitochondrial transport and fatty acid synthesis, cytosolic acetyl-CoA is also a key regulator of histone acetylation and chromatin silencing. It would strengthen the study to consider whether acetyl-CoA depletion contributes to improved fitness through enhanced rDNA silencing. Given the well-established role of rDNA instability in yeast aging, additional experiments examining rDNA silencing and stability would be valuable. For example, monitoring rDNA copy number changes (not necessarily ERCs) under AMPK activation, oleic acid supplementation, and in the A2A strain, similar to approaches used in the authors' prior work, would help clarify whether chromatin regulation contributes to the observed phenotypes.

      We have added data addressing these points to the manuscript and Supplemental Figures 2, 3 and 4, though the outcomes are complex. Histone acetylation changes chromatin accessibility and could therefore alter global gene expression; in accord with this, RNA-seq shows that P<sub>GPD</sub>-SAK1 reduces known age-linked gene expression dysregulation. However, A2A does not further reduce the effect, meaning either that another driver exists in addition to cytosolic acetyl-CoA, or that age-linked gene expression dysregulation is unrelated to cytosolic acetyl-CoA. Oleic acid has little effect on age-linked gene expression dysregulation despite rescuing fitness. With regard to rDNA silencing, transcription of the rDNA intergenic spacer non-coding RNAs promotes ERC formation; we have added data showing that ERC accumulation is not reduced in A2A but slightly higher coherent with the higher replicative age of A2A at 48 h, which suggests silencing is not better in A2A. By RNA-seq, these intergenic spacer transcripts are massively upregulated with age, but this will be a consequence of the increased genomic copy number on ERCs; the upregulation is less in A2A than other conditions, but this arises because the log phase spacer transcript levels are higher and so does not reflect better rDNA silencing. We have previously assayed for heritable changes in rDNA copy number arising during ageing and found (to our surprise) absolutely nothing, so we don't expect any changes under these conditions. The upregulation of transcripts from Sir2-repressed telomeric and MAT loci with age is decreased in P<sub>GPD</sub>-SAK1 and A2A, but the effect size is not different from any other low-expressed genes so we do not think there is a particular effect at loci subject to chromatin silencing (see our previous study Zylstra et al PMID 37643194 for evidence that Sir2-mediated gene silencing is not affected by age). We have added our conclusions from these experiments to the Discussion.

      (2) The current data do not fully distinguish whether AMPK activation and oleic acid supplementation act on distinct subpopulations of aging cells. An alternative explanation is that oleic acid supplementation enhances mitochondrial function and acts additively with AMPK activation, thereby increasing the fraction of cells in the "low senescence" state. Since this distinction is not central to the main conclusions, I suggest softening the language around subpopulation specificity. Emphasizing instead that the A2A strain coordinately modulates multiple branches of acetyl-CoA metabolism to improve late-life fitness would maintain the strength of the central message without over interpretation.

      We respectfully disagree with the reviewer on this point. We show that P<sub>GPD</sub>-SAK1 rescues senescence in ~half the population by a Cat2/Mls1 dependent mechanism (Figure 1F). We then show that in A2A, which rescues most cells, deletion of CAT2/MLS1 restores senescence in ~half the cells (Figure 3F/G). This cannot be explained by an additive mechanism as this would either result in all cells being partially rescued in the P<sub>GPD</sub>-SAK1 and in the A2A cat2Δ mls1Δ mutants, which is definitely not the case either by Tom70-GFP or fitness. Instead the population splits into high/low senescence and fit/unfit cells in the different assays.

      On the specific point of whether lipid synthesis additively increases mitochondrial function, we have added oxygen consumption rate data showing that A2A cells respire more than P<sub>GPD</sub>-SAK1 at 48h but only by a relatively small amount (Figure S3D), so there is indeed an additive improvement in mitochondrial function, but too little to explain the difference in population fitness in our opinion.

      We realise that the reviewer is asking more specifically about oleic acid, but again in the flow data, Figure 4C, what changes with oleic acid or P<sub>GPD</sub>-SAK1 is the proportion of cells in the low Tom70 / high WGA sector. Under an additive effect model, oleic acid or P<sub>GPD</sub>-SAK1 individually would partially reduce Tom70 and partially increase WGA, but the population in the low Tom70 / high WGA sector has the same average Tom70/WGA values in oleic acid, P<sub>GPD</sub>-SAK1 or P<sub>GPD</sub>-SAK1+oleic acid. It is the proportion of cells in this population that changes. Furthermore, under an additive model, wildtype cells aged with oleic acid would not have highest fitness than P<sub>GPD</sub>-SAK1 or A2A (Figure 4D) as these individual cells would lack the mitochondrial upregulation from P<sub>GPD</sub>-SAK1.

      (3) The manuscript proposes that lipid starvation and excess acetyl-CoA are major drivers of senescence in distinct subpopulations of wild-type aging cells. This conclusion is not yet fully supported by the presented data. Direct measurements of age-dependent divergence in acetyl-CoA and fatty acid levels at the single-cell level would be needed to substantiate this model. Based on the current evidence, a more conservative interpretation would be that aging cells exhibit differential sensitivity to perturbations in acetyl-CoA and lipid metabolism. Accordingly, I recommend revising the statement in the Abstract ("We further implicate lipid starvation and excess acetyl coenzyme A availability as major drivers of senescence...") and the corresponding discussion text to better align with the data.

      We agree and have adjusted the abstract to make it clearer that the lipid starvation / excess acetyl-coA interpretation is a model.

      “Our findings support a model in which lipid starvation and excess acetyl-coenzyme A availability are major drivers of senescence in replicatively aged wild-type yeast.”

      Reviewer #3 (Public review):

      Summary:

      These findings suggest that PGPD-SAK1 yeast show a subpopulation with lowered TOM70-GFP expression in high bud scar staining aged cells. Deletion of CAT2 or MLS1 reduces this effect. A PGPD-SAK1 acc1S1157A double mutant (called "A2A" here) shows an even larger effect of lowered tom70 expression in high bud scar staining aged cells. Utilization of various additional mutants involved in acetyl-CoA transport, carnitine shuttle, respiration, etc., leads the authors to conclude that these shifts in TOM70-GFP in aged cells are linked to the AMPK-fatty acid metabolic regulatory system.

      Strengths:

      These extensive and clearly described experiments reveal interesting changes in TOM70-GFP intensity in subsets of aged yeast in several mutants eventually identified as linked to the AMPK-fatty acid metabolic regulatory system.

      Weaknesses:

      (1) 3 biological replicates for mRNASeq is low.

      Thank you for pointing this out. We performed another replicate after posting the initial preprint to confirm the finding but didn’t update the figure in the eLife-reviewed version. We have added this to the scatter plots and analysis in Figure 1, there are minor changes but the set of genes we followed up are still highly significant. For ageing experiments, we sequence to n=3 as a first pass which is sufficient to detect widespread age-linked gene expression effects, and add more replicates if required to solidify findings for specific sets of genes. Hence, the additional RNAseq experiments we have added to the manuscript to Address Reviewer 2’s comments on widespread gene expression effects are also n=3-4.

      (2) While "Traditional conceptions of ageing implicate a progressive accumulation of damage leading to systemic degradation in performance until death, with evolutionary pressures acting to maximise early life fitness and fecundity at the expense of ageing health." is tangential perhaps to the data and conclusions of the study, both claims of this sentence are at best controversial, and the manuscript is no weaker for their omission.

      We would prefer not to remove this sentence, which we see as important to a major message of the manuscript: that ageing does not have to involve a loss of fitness before death. Outside the ageing biology field, ageing is often described as the progressive wearing out of components leading to decline and death (‘like an old car’ is a common analogy); in the ageing field this is certainly controversial, but outside the field it remains the normal understanding. This is what we mean by traditional conceptions, and it is important to consider the contradiction between this widely held viewpoint and our findings (and of course those of many others in the ageing field).

      The second part of the sentence about evolutionary pressures alludes to antagonistic pleiotropy, which we have now made explicit. Antagonistic pleiotropies as a driving mechanism for ageing, while not universally accepted, are as far as we can tell the most widely accepted type of theory in the ageing field. Our interpretation that yeast are bet-hedging as a population growth strategy and this drives ageing in the long term is a classic antagonistic pleiotropy and we need to raise this concept in the introduction.

      (3) The statement that "Here, we determine the basis of senescence and fitness loss in replicatively ageing yeast" is a bit strong as a summary of the present careful work presented here. If the authors had created yeast mutants that retained fitness indefinitely, this would be a more appropriate strength of claim to summarize the work.

      We agree and have moderated this sentence:

      “Here, we show that senescence and fitness loss in replicatively ageing yeast can be almost completely avoided without extension of lifespan by rewiring the conserved AMPK-fatty acid metabolic regulatory system.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The labelling of Figure 3G horizontal axis needs to be realigned with the data.

      Fixed – thank you.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 3G, the x-axis labels appear misaligned and should be corrected for clarity.

      Fixed – thank you.

      (2) Figures S3B and S3C appear to be mislabeled and should be revised.

      Fixed – thank you.

      (3) On page 6 (3rd paragraph), the statement that the beneficial impact arises from acetyl-CoA removal "rather than a benefit of respiration" may be overstated. The data support a role for acetyl-CoA removal but do not fully exclude a contribution from respiration. A more balanced phrasing would improve accuracy.

      We have revised this sentence and also added data:

      “Working in sip2Δ to avoid an increase in AMPK activity due to reduced Acetyl-CoA availability, we observed that ald6Δ increased the low senescence population through decreasing Tom70-GFP (S3C), and therefore the beneficial impact of PGPD-SAK1 on this pathway arises primarily through Acetyl-CoA removal. It is possible that respiration is adding to this benefit, and we detect a significant increase in Oxygen Consumption Rate in aged PGPD-SAK1 cells, but the further increase in A2A is smaller and we consider that this cannot fully explain the effect of acc1S1157A.”

      Reviewer #3 (Recommendations for the authors):

      This manuscript is clearly written, and the data are clearly presented. While 3 biological replicates is inadvisably low for mRNASeq, the subsequent experiments motivated by the genes identified there nevertheless stand on their own as presented.

      Thank you.

    1. eLife Assessment

      This study offers important insight into the pathogenic basis of intragenic frameshift deletions in the carboxy-terminal domain of MECP2, which account for some Rett syndrome cases, yet similar variants also appear in unaffected individuals. Using base editing and mouse models, the authors present convincing evidence supporting the pathogenicity of select deletion variants, with potential implications for therapeutic development.

    2. Reviewer #1 (Public review):

      Summary:

      The authors scrutinized difference in C terminal region variant profiles between Rett syndrome patients and healthy individuals and pinpointed that subtle genetic alternation can cause benign or pathogenic output, which harbors important implication in Rett syndrome diagnosis and proposing therapeutic strategy. This work will be beneficial to clinicians and basic scientists who work on Rett syndrome and carries potential to be applied to other Mendelian rare diseases.

      Strengths:

      Well-designed genetic and molecular experiments, translating genetic differences into functional and clinical changes. This is a unique study resolving subtle changes in sequences give rise to dramatic phenotypic consequences.

      Comments on revised version.

      Improvements were made during the revision.

    3. Reviewer #2 (Public review):

      Summary:

      This study by Guy and Bird and colleagues is a natural follow-up to their 2018 Human Molecular Genetics paper, further clarifying the molecular basis of C-terminal deletions (CTDs) in MECP2 and how they contribute to Rett syndrome. The authors combine human genetic data with well-designed experiments in embryonic stem cells, differentiated neurons, and knock-in mice to explain why some CTD mutations are disease-causing while others are harmless. They show that pathogenic mutations create a specific amino acid motif at the C-terminus, where +2 frameshifts produce a PPX ending that greatly reduces MeCP2 protein levels (likely due to translational stalling) whereas +1 frameshifts generating SPRTX endings are well tolerated.

      Strengths:

      This is a comprehensive and rigorous study that convincingly pinpoints the molecular mechanism behind CTD pathogenicity, with strong agreement between the cell-based and animal data. The authors also provide a proof of principle that modifying the PPX termination codon can restore MeCP2-CTD protein levels and rescue symptoms in mice. In addition, they demonstrate that adenine base editing can correct this defect in cultured cells and increase MeCP2-CTD protein levels. Overall, this is a well-executed study that provides important mechanistic and translational insight into a clinically important class of MECP2 mutations.

      Weaknesses:

      The adenine base editing to change the termination codon is shown feasible in generated cell lines, but yet to be shown in vivo in animal models.

      Comments on revised version.

      The authors have addressed all of my questions and comments.

    4. Author response:

      The following is the authors’ response to the original reviews.

      This is a summary of the changes that have been made to the Reviewed Preprint:

      (1) The data from RettBASE which was analysed in the manuscript has been added in the form of four supplementary tables. Supplementary Table 1 contains the download of all MeCP2 mutations contained in RettBASE. Supplementary Tables 2-4 contain subsets of this data which were used in Figure 2B and Figure 3 SF2. Supplementary Table 4 also has the HGVS nomenclature for both e1 and e2 isoforms and the ClinVar Variation ID for each allele. Wording has been changed to clarify that analysis in the manuscript used this data from RettBASE and not the information that was deposited in ClinVar.

      (2) Similarly, Supplementary Tables 5-8 contain the gnomAD data that was used in the preparation of Figure 2, Figure 2 SF1 and Figure 3 SF1. The “high confidence” alleles in Supplementary Table 8 have been annotated with their HGVS names.

      (3) The criteria for selecting “high confidence” RettBASE and gnomAD alleles have been more explicitly stated in both the Results and Materials and Methods sections.

      (4) An additional “high confidence” RTT mutation (c.1152_1195) has been added to Figures 2B and 3B.

      (5) Figure 3 Supplementary Figure 2 has been added to show reading frame data for all frameshifting deletions in the C-terminal deletion-prone region (CT-DPR), showing that +2 frameshifts predominate in this larger data set, not just in the “high confidence” set. This has necessitated changing the previous Fig. 3 SF2 to Fig. 3 SF3.

      (6) A summary of the genetic alterations described in the manuscript, and their outcomes, has been added as Figure 7.

      (7) A simple flow chart which assists in the classification of human CTDs as “likely benign” or “likely pathogenic” has been added as Figure 8. This will aid future assessment of novel mutations in this region.

      (8) Additions have been made to the Materials and Methods section to comply with reporting guidelines.

      (9) Minor changes have been made to the text to correct typographical errors and to clarify meaning.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors scrutinized differences in C-terminal region variant profiles between Rett syndrome patients and healthy individuals and pinpointed that subtle genetic alternation can cause benign or pathogenic output, which harbors important implications in Rett syndrome diagnosis and proposes a therapeutic strategy. This work will be beneficial to clinicians and basic scientists who work on Rett syndrome, and carries the potential to be applied to other Mendelian rare diseases.

      Strengths:

      Well-designed genetic and molecular experiments translate genetic differences into functional and clinical changes. This is a unique study resolving subtle changes in sequences that give rise to dramatic phenotypic consequences.

      Weaknesses:

      There are many base-editing and protein-expression changes throughout the manuscript, and they cause confusion. It would be helpful to readers if authors could provide a simple summary diagram at the end of the paper.

      We have added a summary diagram, as suggested (Figure 7). We have also provided a flowchart which shows how to classify human CTDs as “likely benign” or “likely pathogenic” based on their location.

      Reviewer #2 (Public review):

      Summary:

      This study by Guy and Bird and colleagues is a natural follow-up to their 2018 Human Molecular Genetics paper, further clarifying the molecular basis of C-terminal deletions (CTDs) in MECP2 and how they contribute to Rett syndrome. The authors combine human genetic data with well-designed experiments in embryonic stem cells, differentiated neurons, and knock-in mice to explain why some CTD mutations are disease-causing while others are harmless. They show that pathogenic mutations create a specific amino acid motif at the C-terminus, where +2 frameshifts produce a PPX ending that greatly reduces MeCP2 protein levels (likely due to translational stalling) whereas +1 frameshifts generating SPRTX endings are well tolerated.

      Strengths:

      This is a comprehensive and rigorous study that convincingly pinpoints the molecular mechanism behind CTD pathogenicity, with strong agreement between the cell-based and animal data. The authors also provide a proof of principle that modifying the PPX termination codon can restore MeCP2-CTD protein levels and rescue symptoms in mice. In addition, they demonstrate that adenine base editing can correct this defect in cultured cells and increase MeCP2-CTD protein levels. Overall, this is a well-executed study that provides important mechanistic and translational insight into a clinically important class of MECP2 mutations.

      Weaknesses:

      The adenine base editing to change the termination codon is shown to be feasible in generated cell lines, but has yet to be shown in vivo in animal models.

      This work is the obvious next step and is in progress. However, with the rise in pre- and neonatal genetic testing we felt it was important to disseminate our findings as soon as possible. The family pedigree in Figure 3C is a clear illustration of this need

      Reviewer #3 (Public review):

      Summary:

      Guy et al. explored the variation in the pathogenicity of carboxy-terminal frameshift deletions in the X-linked MECP2 gene. Loss-of-function variants in MECP2 are associated with Rett syndrome, a severe neurodevelopmental disorder. Although 100's of pathogenic MECP2 variants have been found in people with Rett syndrome, 8 recurrent point mutations are found in ~65% of disease cases, and frameshift insertions/deletions (indels) variants resulting in production of carboxy-terminal truncated (CTT) MeCP2 protein account for ~10% of cases. Many of these occur in a "deletion prone region" (DPR) between c.1110-1210, with common recurrent deletions c.1157-1197del (CTD1) and c.1164_1207del (CTD2). While two major protein functional domains have been defined in MeCP2, the methyl-binding domain (MBD) and the NCoR interacting domain (NID), the functional role of the carboxy-terminal domain (CTD, beyond the NID, predicted to have a disordered protein structure) has not been identified, and previous work by this group and others demonstrated that a Mecp2 "minigene" lacking the CTD retains MeCP2 function suggesting that the CTD is dispensable. This raises an important question: If the CTD is dispensable, what is the pathogenic basis of the various CTT frameshift variants? Prior work from this group demonstrated that genetically engineered mice expressing the CTD1 variant had decreased expression of Mecp2 RNA and MeCP2 protein and decreased survival, but those expressing the CTD2 variant had normal Mecp2 RNA and protein and survival. However, they noted that differences between the mouse and human coding sequences resulted in different terminal sequences between the two common CTD, with CTD1 ending in -PPX in both mouse and human, but CTD2 ending in -PPC in human but -SPX in mouse, and in the previous paper they demonstrated in humanized mouse ES cells (edited to have the same -PPX termination) containing the CTD2 deletion resulted in decreased Mecp2 RNA and protein levels. This previous work provides the underlying hypotheses that they sought to explore, which is that the pathological basis of disease causing CTD relates to the formation of truncated proteins that end with a specific amino acid sequence (-PPX), which leads to decreased mRNA and protein levels, whereas tolerated, non-pathogenic CTD do not lead to production of truncated proteins ending in this sequence and retain normal mRNA/protein expression.

      In this manuscript, they evaluate missense variants, in-frame deletions, and frame shift deletions within the DPR from the aggregated Genome Aggregated Database (gnomAD) and find that the "apparently" normal individuals within gnomAD had numerous tolerated missense variants and in-frame deletions within this region, as well as frameshift deletions (in hemizygous males) in the defined region. All of the gnomAD deletions within this region resulted in terminal amino acid sequences -SPRTX (due to +1 frameshift), whereas nearly all deletion variants in this region from people with Rett syndrome (from the Clinvar copy of the former RettBase database) had a terminal -PPX sequence, due to a +2 frameshift. They hypothesized that terminal proline codons causing ribosomal stalling and "nonsense mediated decay like" degradation of mRNA (with subsequent decreased protein expression) was the basis of the specific pathogenicity of the +2 frameshift variants, and that utilizing adenine base editors (ABE) to convert the termination codon to a tryptophan could correct this issue. They demonstrate this by engineering the change into mouse embryonic stem cell lines and mouse lines containing the CTD1 deletion and show that this change normalized Mecp2 mRNA and protein levels and mouse phenotypes. Finally, they performed an initial proof-of-concept in an inducible HEK cell line and showed the ability of targeted ABE to edit the correct adenine and cause production of the expected larger truncated Mecp2 protein from CTD1 constructs.

      The findings of this manuscript provide a level of support for their hypothesis about the pathogenicity versus non-pathogenicity of some MECP2 CTT intragenic deletions and provide preliminary evidence for a novel therapeutic approach for Rett syndrome; however, limitations in their analysis do not fully support the broader conclusions presented.

      Strengths:

      (1) Utilization of publicly available databases containing aggregated genetic sequencing data from adult cohorts (gnomAD) and people with Rett syndrome (Clinvar copy of RettBase) to compare differences in the composition of the resulting terminal amino acid sequences resulting from deletions presumed to be pathogenic (n+2) versus presumed to be tolerated (n+1).

      (2) Evaluation of a unique human pedigree containing an n+1 deletion in this region that was reported as pathogenic, with demonstration of inheritance of this from the unaffected father and presence within other unaffected family members.

      (3) Development of a novel engineered mouse model of a previously assumed n+1 pathogenic variant to demonstrate lack of detrimental effect, supporting that this is likely a benign variant and not causative of Rett syndrome.

      (4) Creation and evaluation of novel cell lines and mouse models to test the hypothesis that the pathogenicity of the n+2 deletion variants could be altered by a single base change in the frameshifted stop codon.

      (5) Initial proof-of-concept experiments demonstrating the potential of ABE to correct the pathogenicity of these n+2 deletion variants.

      Weaknesses:

      (1) While the use of the large aggregated gnomAD genetic data benefits from the overall size of the data, the presence of genetic variants within this collection does not inherently mean that they are "neutral" or benign. While gnomAD does not include children, it does include aggregated data from a variety of projects targeting neuropsychiatric (and other conditions), so there is information in gnomAD from people with various medical/neuropsychiatric conditions. The authors do make some acknowledgement of this and argue that the presence of intragenic deletion variants in their region of interest in hemizygous males indicates that it is highly likely that these are tolerated, non-pathogenic variants. Broadly, it is likely true that gnomAD MECP2 variants found in hemizygous males are unlikely to cause Rett syndrome in heterozygous females, it does not necessarily mean that these variants have no potential to cause other, milder, neuropsychiatric disorders. As a clear example, within gnomAD, there is a hemizygous male with the rs28934908 C>T variant that results in p.A140V (p.A152V in e1 transcript numbering convention). This pathogenic variant has been found in a number of pedigrees with an X-linked intellectual disability pattern, in which males have a clear neurodevelopmental disorder and heterozygous females have mild intellectual disability (see PMIDs 12325019, 24328834 as representative examples of a large number of publications describing this). Thus, while their claim that hemizygous deletion variants in gnomAD are unlikely to cause Rett syndrome, that cannot make the definitive statement that they are not pathogenic and completely benign, especially when only found in a very small number of individuals in gnomAD.

      We have included the possibility that mutations found in gnomAD may give rise to less severe neurological conditions in the discussion.

      (2) The authors focus exclusively on deletions within the "DPR", they define as between c.1110-1210 and say that these deletions account for 10% of Rett syndrome cases. However, the published studies that are the basis for this 10% estimate include all genetic variants (frameshift deletions, insertions, complex insertion/deletions, nonsense variants) resulting in truncations beyond the NID. For example, Bebbington 2010 (PMID: 19914908), which includes frameshift indels as early as c.905 and beyond c.1210. Further specific examples from RettBase are described below, but the important point is that their evaluation of only frameshift variants within c.1110-1210 is not truly representative of the totality of genetic variants that collectively are considered CTT and account for 10% of Rett cases.

      The vast majority of C-terminal truncating mutations do occur within the “CT-DPR”, likely due to its C-rich nucleotide sequence and the presence of microhomologies within the region. Looking at frameshifting deletions in RettBASE that start after the NID, a large proportion of these end within the CT-DPR and result in a -PPX ending. We decided to restrict our analysis to the c.1110-1210 region to avoid including the rarer examples that may have a different reason for their pathogenicity. We do not assert that all C-terminal truncations are pathogenic due to this mechanism, but current evidence suggests that most are.

      (3) The authors say that they evaluated the putative pathogenic variants contained within RettBase (which is no longer available, but the data were transferred to Clinvar) for all cases with Classic Rett syndrome and de novo deletion variants within their defined DPR domain. Looking at the data from the Clinvar copy of RettBase, there are a number (n=143) of c-terminal truncating variants (either frameshift or nonsense) present beyond the NID, but the authors only discuss 14 deletion frameshift variants in this manuscript. A number of these variants have molecular features that do not fall into the pathogenic classification proposed by the authors and are not addressed in the manuscript and do not support the generalization of the conclusions presented in this manuscript, especially the conclusion that the determination of pathogenicity of all c-terminal truncating variants can be determined according to their proposed n+2 rule, or that all of the 10% of people with Rett syndrome and c-terminal truncating variants could be treated by using a base editor to correct the -PPX termination codon.

      It is important to state here that we did not use the data in ClinVar for our analysis, but the original information that was held in RettBASE. We have clarified this in the manuscript and have now included supplementary tables containing the data we downloaded from both RettBASE and gnomAD. Table 1 contains all the RettBASE entries with MECP2 mutations, while Tables 2-4 contain subsets of this data pertaining to CTDs. We have extended our “high confidence” set of RTT alleles to contain one more that was previously overlooked (c.1152_1195) to bring the total number of alleles to 15. Taken together these alleles account for 158 individual entries in RettBASE. We have now included an analysis of all frameshifting mutations in the CT-DPR (Supplementary Table 3, Fig. 3 SF2) which covers 69 different mutations and 260 individual entries. Of these, +2 frameshifts make up the large majority, in contrast to the gnomAD data shown in Supplementary Table 8 and Fig. 3 SF1.

      (4) The HEK-based system utilized is convenient for doing the initial experiments testing ABE; however, it represents an artificial system expressing cDNA without splicing. Canonical NMD is dependent on splicing, and while non-canonical "NMD-like" processes are less well understood, a concern is whether the artificial system used can adequately predict efficacy in a native setting that includes introns and splicing.

      We disagree with this opinion. We show that the loss of protein and mRNA seen with knock in mouse and human alleles is recapitulated when using a cDNA-based transgene in the HEK system, demonstrating that the mechanism of loss does not involve factors bound at splice junctions etc. We also demonstrate the effect of the A to G change at the stop codon is the same whether we do this by base editing our cDNA transgene in T-REx cells or by making the CTD1 X>W knock-in mouse. Both result in increased levels of a slightly extended but still truncated protein.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The phrase in the title, "an alternative therapeutic approach" is only insinuated in the manuscript, making it rather inappropriate to be in the title.

      In this study we use adenine base editing to modify RTT-causing CTD mutations in T-REx cells which is clearly a precursor to developing a therapy, utilising the new findings in this study. We therefore feel that the use of “an alternative therapeutic approach” can be justified.

      Reviewer #2 (Recommendations for the authors):

      I have a few minor comments for the authors to consider:

      (1) Please double-check Figure 2 Supplementary Figure 1, as the allele count for E394K does not appear to be in the thousands; rather, E397K seems to be the variant shown in the graph.

      Yes, this was an error and has been corrected to E397K in the text. Thank you for spotting it.

      (2) On page 12, the phrase “common DNA sequence features shared by all CTDs that give rise to RTT” might be better described as “amino acid sequence features.”

      This has been altered in the text as suggested.

      (3) On page 3, the sentence "analysis of patient mutations and experimental data from mouse models support a role in transcriptional repression" cites Gabel et al. 2015 and Kinde et al. 2016, which focus on null alleles but not patient mutations. It would be appropriate to also cite Johnson et al. 2017, which analyzed MeCP2 T158M and R106W patient mutations.

      This has now been cited.

      (4) In the same sentence, Bajikar et al. 2025 are described as studying the "acute loss of MeCP2," but gene expression was analyzed after a week or longer period of time, not minutes to hours as in degron-mediated degradation systems.

      The word “acute” has been removed from the text.

      (5) On page 7, the statement that "E394K is common... who were later found to have additional pathogenic MECP2 mutations" should include a supporting reference.

      We have now included the references Moncla et al (2002) and Wan et al (1999) to address this.

      (6) Similarly, on page 10, the sentence "these findings question the validity of two cases where individuals presented with classical Rett..." is missing a reference to the case report mentioned.

      The references Bienvenu et al (2000) and Philippe et al (2006) have been added.

      (7) While the manuscript is well written and full of detail, if space is an issue, the authors might consider tightening sections that reiterate findings from their 2018 HMG paper.

      A section discussing the CTD2 allele from the 2018 HMG paper has been removed from the results section.

      Reviewer #3 (Recommendations for the authors):

      (1) Overall, the manuscript is rather dense and potentially challenging to follow easily, especially for a non-expert reader.

      Minor edits have been made to the text which will hopefully make it easier to follow. We have also added two new figures (Figures 7 and 8) to summarize the different alleles and edits which appear in the paper, and to show how to determine whether a CTD in the region is likely to be benign or pathogenic.

      (2) The introduction of data presented in Figure 1 within the manuscript introduction seems inappropriate and should be moved to the results section.

      We would say that this is unconventional rather than inappropriate, and is referred to in the introduction, so we would prefer to leave it as it is.

      (3) Providing specific, common nomenclature for genetic repository variants (rs numbers, gnomAD IDs, etc) somewhere would be beneficial. This is an issue because of the complexity of numbering (either coding or protein) for MECP2 due to the different transcript-based numbering systems.

      This nomenclature is now included in the supplementary tables of data from RettBASE and gnomAD.

      (4) As described in the public comments, there are a number of MECP2 genetic variants listed in the Clinvar copy of RettBase, resulting in c-terminal truncations that are not mentioned or discussed within the manuscript. Without the level of detail present in the original form of RettBase (number of events, de novo, etc) in the currently available Clinvar iteration, it is unclear why a number of variants, even within the limited DPR region, were not mentioned. A supplementary file including the more complete information from the RettBase version, with a complete listing of all c-terminal truncating variants, and an explicit rationale for the exclusion of variants would be helpful.

      We have now included supplementary tables with our download of all MECP2 mutations which were held in RettBASE. We have further added tables with the subset of mutations that we have analysed and have more explicitly stated our criteria for defining the “high confidence” sets of mutations. We did not download the data relating to “evidence of pathogenicity” (ie de novo?, absent from parents etc) from RettBASE, but annotated our list of CTDs with this information while RettBASE was still available. This was used in Supplementary Table 3.

      (5) A discussion of the limitations, notably that the fact that the focus exclusively on deletion variants within a restricted region (c.1110-1210) does not truly represent all genetic variants that cumulatively account for 10% of Rett cases, is needed. Furthermore, as pointed out, not all frameshift variants, even those that are n+2, result in the -PPX termination that is presented as the pathogenic basis of c-terminal truncations and amenable to correction by ABE. This should be noted in the discussion, as well as consideration of the late nonsense variants that cause c-terminal truncations (some of which would be very similar to the deletion variants discussed but without the proposed primary pathogenic driver, -PPX).

      We do not claim to explain the pathogenic mechanism of all C-terminal frameshift mutations found in cases of Rett syndrome. There will certainly be some that do not fit our explanation. However, we believe we have shown evidence that a large proportion of CTDs in RTT will be amenable to the therapy we propose.

      (6) Regarding point 3 in the public review, specifically:

      (a) n=7 nonsense variants (S360X, K363X, E397X, R453X, E455X) that do not carry the destabilizing -PPX sequence.

      (b) n=136 frameshift indel variants beyond the NID.

      (i) n=11 that have indels that extend past the native stop codon, n=4 of which start within the DPR domain (c.1110-1210) but would have a different terminal sequence than their proposed pathogenic -PPX sequence.

      (ii) n=125 frameshift indels with terminal breakpoint before the native stop codon

      (c) n=89 that have start or stop points within c.1110-1210

      (d) n=72 not mentioned within the manuscript.

      (e) n=27 are n+1, with 26/27 having what the authors term as the "tolerated" -SPRTX ending, but 1/27 having a frameshift beyond this region (c.1133_1361)

      (f) n=45 are n+2, with 32/45 ending in -PPX (supporting authors conclusion), but 10/45 will use the frameshift stop codon preceding the -PPX and have a different terminal sequence, and 3/45 result in a frameshift termination beyond the -PPX sequence.

      (g) n=36 have breakpoints either before c.1110 or after c.1210

      (h) n=16 start before c.1110, with 11/16 ending before c.1110. 4 of these 11 are n+2, but would use the earlier frameshift stop codon and not have -PPX terminal sequence. For 5/16, the indel extends past c.1210, with the n+2 leading to frameshift termination codons beyond the -PPX sequence.

      (i) n=20 indels start beyond c.1210, with 8/20 being n+1 and 12/20 being n+2, with neither leading to the -SPRTX or --PPX termination sequences characterized in this manuscript.

      As mentioned previously, we have used data taken from RettBASE, not from ClinVar. Both RettBASE and ClinVar will contain MECP2 mutations found in cases of RTT which are not the causative mutation. Databases of this kind contain sequencing errors and mutations that have been mistakenly assigned as causative. It is therefore imperative that the publicly available information is screened to only include mutations that meet stringent criteria. This is why we chose to start by looking at high confidence sets of mutations, with our conclusions supported by analysis of all such mutations in our region of study.

      As mentioned in response to point 5 above, we do not claim to explain every mutation in the region, but believe this study reveals an important disease mechanism for a large proportion of CTDs, leading to a potential therapy. It also contains significant information for predicting the likely prognosis of individuals with CTDs, who may remain healthy but are currently informed that their mutation is likely pathogenic. At present it seems that this is often based solely on the presence of a frameshifting mutation with similarities to bona fide RTT CTDs, without strong evidence of pathogenicity.

    1. eLife Assessment

      This valuable study presents a comprehensive exploration of c-di-GMP-associated protein interaction networks in Pseudomonas fluorescens, with a particular focus on biofilm-related phenotypes. The evidence is convincing, supported by a genome-wide yeast two-hybrid screen, phenotypic analyses, and experimental validation. The work identifies multiple interaction hubs and provides a resource that will be of use to the biofilm and c-di-GMP communities, while additional mechanistic exploration would further enhance its impact by clarifying the biological roles of many of the identified interactions.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript by Noirot-Gros et. al. presents a herculean effort to map the protein-protein interactome of the c-di-GMP signaling network in Pseudomonas fluorescens (Pf). C-di-GMP, the key driver of biofilm formation in bacteria, is controlled by a highly complex network of synthesis, degradation and effector proteins. Pf is no exception as it encodes dozens of such proteins. The authors use a Yeast Two-Hybrid approach genome-wide screen with 10 diguanylate cyclase (DGC) enzymes as bait to assess protein-protein interactions in this network. The results identify over one hundred such interactions with several different hubs, including c-di-GMP signaling, other signaling systems, membrane proteins, etc. The authors then explore the original bait proteins as well as identify interactors on biofilm formation-related phenotypes and swarming using a high-throughput CRISPRi expression knockdown approach. The amount of data generated is quite impressive. Much of the manuscript uses statistical-based network analysis to group different proteins based on their interactions or impact on phenotypes, which is a high-level analysis that can catalyze further study into this system. The authors chose three specific proteins to assess their impact on cell morphology, DNA repair, and protein localization. Overall, in my view, this is perhaps the best analysis of a c-di-GMP protein-protein interactome, and it provides a multitude of hypotheses to be tested. However, therein lies the weakness of the manuscript in that very few of these hypotheses are actually tested. But such is not the goal of this network analysis type of approach. Overall, I think the work will be highly impactful to those in the c-di-GMP field, and it provides a template for others attempting such analyses of protein-protein interactions.

      Strengths:

      The manuscript is impressive in the sheer scale of the protein-protein interactions identified, network analysis, and phenotypic analysis of specific proteins in the network. It is an impressive amount of work that could be very useful to the field. It is also statistically rigorous in its analysis of significant interactions or network nodes.

      Weaknesses:

      The weakness of the manuscript is that, with three exceptions, very few of the hypotheses are actually tested. For example, BifA is shown to be a network hub protein that interacts with many other diguanylate cyclases, and this is hypothesized to be through GGDEF heterodimerization. I appreciate that experimentally testing such a hypothesis is probably another entire manuscript, but some early forays into such ideas could be undertaken using AlphaFold structural modeling of protein-protein interactions compared with GGDEFs that don't form heterodimers. Also, an inherent weakness is that such detailed analyses of a c-di-GMP signaling network, in which each diguanylate cyclase and phosphodiesterase may respond to a unique cue, is that the network identified and the conclusions made are highly specific to the experimental conditions in which the work was done. Therefore, it is unclear how broadly these conclusions (i.e. BifA is the central regulator of c-di-GMP signaling) apply to other conditions. But it is impossible to get around such a limitation, and this work can lead to testing the robustness of the identified network in other environments.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Noirot-Gros and coworkers investigated the network of c-di-GMP associated protein complexes in Pseudomonas fluorescens. They did so by using a genome-wide yeast two-hybrid screen, and that was further probed by phenotypic screening that focused on biofilm and motility phenotypes. From this network map, they discovered that the phosphodiesterase DipA interacts with the GGDEF domains of many c-di-GMP-binding proteins.

      Strengths:

      (1) Broadness of screen led to identification of new interactions: The genome-wide yeast two-hybrid screening approach permitted broad investigation of c-di-GMP-associated protein-protein interactions. These interactions included some previously validated interactions as well as newly discovered interactions.

      (2) Complementary experimental validation: The proposed network was experimentally validated, including by using a CRISPRi-based approach in which the expression of genes encoding proteins identified in the network was systematically suppressed, and then the impact on the biofilm and motility phenotypes was assessed.

      Weaknesses:

      The findings would have been strengthened by further biochemical analysis, but this is likely beyond the scope of the paper.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Noirot-Gross et al take an open-ended approach to elucidate the c-diGMP-associated protein complexes in Pseudomonas fluorescens. Starting with 10 cyclic d-GMP putative proteins, they use a combination of genome-wide two-hybrid system followed by CRISPRi-mediated exploration of phenotypes to describe the cyclic di-GMP-associated regulation of biofilm formation, and how it relates to other functions. Overall, this work presents an excellent example of how genome annotations can be further confirmed with the use of integrated functional genomic approaches. Some areas of improvement can be applied to this manuscript to enhance readability and provide a clearer distinction between confirmatory results and new findings, which are provided below:

      Strengths:

      (1) The authors have explored their findings extensively and provide a comprehensive view of the topic.

      (2) The combination of genome-wide explorations of protein-protein interactions with the more focused phenotypic exploration of the interactions found provides a solid framework for the work presented.

      Weaknesses:

      (1) Overall goal of the work:

      While articles that describe open-ended approaches can be comprehensive and descriptive in nature, the authors should have a main overall goal, which can guide the reader through the main and most compelling findings at the end. As written, the overall goal is not clear. The network perspective is interesting, and the focus on biofilm formation appears in the title. Why P. fluorescens? How is cyclic di-GMP-mediated regulation of biofilm formation in P. fluorescens different from P. aeruginosa? Why would it be studied?

      (Positive or negative regulation of biofilm formation?)

      (2) Abstract:

      The abstract is very well written and guides the reader to the DipA as a hub protein in the network. From further reading, the article could clarify whether this finding is confirmatory or novel (does DipA play a similar role in P. aeruginosa?) It would be appropriate to mention the role of DipA in other Pseudomonas species from the beginning, and not only in the discussion session.

      (3) Introduction:

      The introduction is nicely written. An area of improvement could be giving more attention to protein interactions as relevant to c-di-GMP. The authors could consider an independent paragraph starting with line 84-85 "Protein-protein interactions involving DGCs, PDEs, and target effectors are crucial in establishing localized signalling through the generation of local pools of c-di-GMP", expanding on this particular aspect with an example of localized signal, after explaining that localization could help decipher specific function within the network of DGCs and PDEs. Then go into connecting biofilms with c-di-GMP and protein-protein interactions, using the example of GcbC and LapD.

      (4) The rationale of choosing 10 PDEs could be clarified. The nice diagrams shown in the supplementary table could be used as part of Figure 1, so the reader understands why these proteins were used, and what is known about them (for example, add them as Figure 1a).

      (5) Figures 1b and 2 convey the same information as in Figure 1a. They could be removed without affecting the understanding of the article.

      (6) CRISPRi and Figure 3. Figure 3 shows the methodology of CRICPR phenotypic screening. A diagram showing the CRISPRi system in P. fluorescens could help the non-expert reader. While the choice of 23 proteins related to the emerging hub DipA is clear, the choice of the other 33 genes could be better explained. Are these proteins already related to biofilm formation? Where are they part of the network detected? How about the other 14 SBW25 genes? The authors could clarify the rationale of the choices. Figure 4 could be combined with Figure 3 or moved to the supplementary material.

      (7) Figures 5, 6 and 7 represent solid network analysis of the findings. Still, they could be improved in clarity on the main findings. The authors conclude at the end of section 3.2.3 that there are networks that exert a "positive role" and a "negative role". The authors could show that in the figures, explaining what those roles are: more biofilm structural coding genes? positive or negative regulation of biofilm formation?)

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Noirot-Gros et. al. presents a herculean effort to map the protein-protein interactome of the c-di-GMP signaling network in Pseudomonas fluorescens (Pf). C-di-GMP, the key driver of biofilm formation in bacteria, is controlled by a highly complex network of synthesis, degradation and effector proteins. Pf is no exception as it encodes dozens of such proteins. The authors use a Yeast Two-Hybrid approach genome-wide screen with 10 diguanylate cyclase (DGC) enzymes as bait to assess protein-protein interactions in this network. The results identify over one hundred such interactions with several different hubs, including c-di-GMP signaling, other signaling systems, membrane proteins, etc. The authors then explore the original bait proteins as well as identify interactors on biofilm formation-related phenotypes and swarming using a high-throughput CRISPRi expression knockdown approach. The amount of data generated is quite impressive. Much of the manuscript uses statistical-based network analysis to group different proteins based on their interactions or impact on phenotypes, which is a high-level analysis that can catalyze further study into this system. The authors chose three specific proteins to assess their impact on cell morphology, DNA repair, and protein localization. Overall, in my view, this is perhaps the best analysis of a c-di-GMP protein-protein interactome, and it provides a multitude of hypotheses to be tested. However, therein lies the weakness of the manuscript in that very few of these hypotheses are actually tested. But such is not the goal of this network analysis type of approach. Overall, I think the work will be highly impactful to those in the c-di-GMP field, and it provides a template for others attempting such analyses of protein-protein interactions.

      Strengths:

      The manuscript is impressive in the sheer scale of the protein-protein interactions identified, network analysis, and phenotypic analysis of specific proteins in the network. It is an impressive amount of work that could be very useful to the field. It is also statistically rigorous in its analysis of significant interactions or network nodes.

      Weaknesses:

      The weakness of the manuscript is that, with three exceptions, very few of the hypotheses are actually tested. For example, BifA is shown to be a network hub protein that interacts with many other diguanylate cyclases, and this is hypothesized to be through GGDEF heterodimerization. I appreciate that experimentally testing such a hypothesis is probably another entire manuscript, but some early forays into such ideas could be undertaken using AlphaFold structural modeling of protein-protein interactions compared with GGDEFs that don't form heterodimers. Also, an inherent weakness is that such detailed analyses of a c-di-GMP signaling network, in which each diguanylate cyclase and phosphodiesterase may respond to a unique cue, is that the network identified and the conclusions made are highly specific to the experimental conditions in which the work was done. Therefore, it is unclear how broadly these conclusions (i.e. BifA is the central regulator of c-di-GMP signaling) apply to other conditions. But it is impossible to get around such a limitation, and this work can lead to testing the robustness of the identified network in other environments.

      We would like to thank the reviewer sincerely for their positive comments on our manuscript and for their constructive feedback. We recognize the limitations arising from the lack of extensive knowledge regarding the environmental cues that trigger the regulation of all CDG activities in P. fluorescens. We hypothesize that DipA acts as a central local hub that positively or negatively regulates the activity of its interacting CDG partners throughout the cell life cycle, lifestyle transitions and environmental signals. Testing this hypothesis would indeed require extensive biochemical and omics approaches. However, strengthening the significance of DipA complexes in silico using AlphaFold is a very appealing proposition and we are currently considering including this analysis in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Noirot-Gros and coworkers investigated the network of c-di-GMP associated protein complexes in Pseudomonas fluorescens. They did so by using a genome-wide yeast two-hybrid screen, and that was further probed by phenotypic screening that focused on biofilm and motility phenotypes. From this network map, they discovered that the phosphodiesterase DipA interacts with the GGDEF domains of many c-di-GMP-binding proteins.

      Strengths:

      (1) Broadness of screen led to identification of new interactions: The genome-wide yeast two-hybrid screening approach permitted broad investigation of c-di-GMP-associated protein-protein interactions. These interactions included some previously validated interactions as well as newly discovered interactions.

      (2) Complementary experimental validation: The proposed network was experimentally validated, including by using a CRISPRi-based approach in which the expression of genes encoding proteins identified in the network was systematically suppressed, and then the impact on the biofilm and motility phenotypes was assessed.

      Weaknesses:

      The findings would have been strengthened by further biochemical analysis, but this is likely beyond the scope of the paper.

      We would like to express our gratitude to the reviewer for their positive evaluation assessment, and for taking into account the limitations of the study's scope.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Noirot-Gross et al take an open-ended approach to elucidate the c-diGMP-associated protein complexes in Pseudomonas fluorescens. Starting with 10 cyclic d-GMP putative proteins, they use a combination of genome-wide two-hybrid system followed by CRISPRi-mediated exploration of phenotypes to describe the cyclic di-GMP-associated regulation of biofilm formation, and how it relates to other functions. Overall, this work presents an excellent example of how genome annotations can be further confirmed with the use of integrated functional genomic approaches. Some areas of improvement can be applied to this manuscript to enhance readability and provide a clearer distinction between confirmatory results and new findings, which are provided below:

      Strengths:

      (1) The authors have explored their findings extensively and provide a comprehensive view of the topic.

      (2) The combination of genome-wide explorations of protein-protein interactions with the more focused phenotypic exploration of the interactions found provides a solid framework for the work presented.

      Weaknesses:

      (1) Overall goal of the work:

      While articles that describe open-ended approaches can be comprehensive and descriptive in nature, the authors should have a main overall goal, which can guide the reader through the main and most compelling findings at the end. As written, the overall goal is not clear. The network perspective is interesting, and the focus on biofilm formation appears in the title. Why P. fluorescens? How is cyclic di-GMP-mediated regulation of biofilm formation in P. fluorescens different from P. aeruginosa? Why would it be studied? (Positive or negative regulation of biofilm formation?)

      We would like to express our appreciation to the reviewer for their thorough evaluation of our manuscript and for the constructive feedback they provided. The overall goal of this study will be further refined, and outlined in the introduction in the revised version of the manuscript.

      (2) Abstract:

      The abstract is very well written and guides the reader to the DipA as a hub protein in the network. From further reading, the article could clarify whether this finding is confirmatory or novel (does DipA play a similar role in P. aeruginosa?) It would be appropriate to mention the role of DipA in other Pseudomonas species from the beginning, and not only in the discussion session.

      (3) Introduction:

      The introduction is nicely written. An area of improvement could be giving more attention to protein interactions as relevant to c-di-GMP. The authors could consider an independent paragraph starting with line 84-85 "Protein-protein interactions involving DGCs, PDEs, and target effectors are crucial in establishing localized signalling through the generation of local pools of c-di-GMP", expanding on this particular aspect with an example of localized signal, after explaining that localization could help decipher specific function within the network of DGCs and PDEs. Then go into connecting biofilms with c-di-GMP and protein-protein interactions, using the example of GcbC and LapD.

      We propose highlighting the example to the local signalling cascade formed by the tripartite system YdaM, YciR and MlrA. This will be addressed in the revised version of the manuscript.

      (4) The rationale of choosing 10 PDEs could be clarified. The nice diagrams shown in the supplementary table could be used as part of Figure 1, so the reader understands why these proteins were used, and what is known about them (for example, add them as Figure 1a).

      We propose to include a specific section in the supplementary file to explain the whole rationale behind choosing these CDGs. These proteins were selected based on their involvement in different steps of biofilm formation in Pseudomonas, as well as their role in the ability of P. fluorescens strains to colonize plant roots.

      (5) Figures 1b and 2 convey the same information as in Figure 1a. They could be removed without affecting the understanding of the article.

      Figure 2 will be transferred in Supplementary as part of the Figure S1

      (6) CRISPRi and Figure 3. Figure 3 shows the methodology of CRICPR phenotypic screening. A diagram showing the CRISPRi system in P. fluorescens could help the non-expert reader. While the choice of 23 proteins related to the emerging hub DipA is clear, the choice of the other 33 genes could be better explained. Are these proteins already related to biofilm formation? Where are they part of the network detected? How about the other 14 SBW25 genes? The authors could clarify the rationale of the choices. Figure 4 could be combined with Figure 3 or moved to the supplementary material.

      A better description of the rationale behind the choice of tested interacting protein partners will be provided. We also agree to combine Figure 4 with Figure 3.

      (7) Figures 5, 6 and 7 represent solid network analysis of the findings. Still, they could be improved in clarity on the main findings. The authors conclude at the end of section 3.2.3 that there are networks that exert a "positive role" and a "negative role". The authors could show that in the figures, explaining what those roles are: more biofilm structural coding genes? positive or negative regulation of biofilm formation?)

    1. eLife Assessment

      This important study conducted by Hall and colleagues advances our understanding of host-pathogen co-evolution by providing an elegant multi-omic comparative framework that uncovers a profound, pathogen-directed epigenomic reprogramming of bovine alveolar macrophages uniquely driven by host-adapted Mycobacterium bovis. Convincing evidence supports these findings and draws on robust, multi-layered high-throughput sequencing data integrated with large-scale cattle GWAS metadata to prioritize actionable genetic variants linked to disease susceptibility and to highlight key immune response genes/pathways upregulated in response to infection and pathogen-specific host adaptation mechanisms.

    2. Reviewer #1 (Public review):

      This manuscript by Hall et al. uses a multi-omic approach to investigate how distinct members of the Mycobacterium tuberculosis complex (MTBC: M. bovis, M. tuberculosis, the attenuated M. bovis BCG vaccine strain, and gamma-irradiated M. bovis) affect bovine alveolar macrophage epigenetic and transcriptional responses after 24 hours. The investigators used RNA-Seq, ATAC-Seq, and ChIP-Seq to assess differential gene expression in each complex type and integrated gene transcription with chromatin accessibility and epigenetic modifications, highlighting key immune response genes/pathways upregulated in response to infection and pathogen-specific host adaptation mechanisms. The analysis also revealed that the most pronounced transcriptional and epigenetic responses were in the M. bovis-infected cells compared with the other complex types. Comparing top genes associated with M. bovis infection of macrophages to a GWAS data set revealed 4 key genes associated with increased susceptibility to infection.

      Overall, this is a technically sound manuscript that contains highly interesting and useful data on bovine innate immune responses to different types of Mycobacterium tuberculosis, which are important to the immunology and infectious disease community as well as the livestock industry. However, in its current format, the manuscript presents the data/figures in a way that is not particularly informative (despite the rich data set) and is too descriptive. We also have some general concerns and suggestions listed below.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Hall et al. present a rigorous and comprehensive multi-omic comparative analysis investigating how host-adapted and non-host-adapted mycobacteria reconfigure the bovine host immune response. By utilizing RNA-seq, ATAC-seq, and ChIP-seq across four histone marks (H3K4me3, H3K4me1, H3K27ac, and H3K27me3) plus CTCF binding, the authors track the regulatory dynamics of primary bovine alveolar macrophages (bAM) challenged with Mycobacterium bovis (MBO), Mycobacterium tuberculosis (MTU), M. bovis BCG, and gamma-irradiated (killed) M. bovis (IRR).

      The study highlights a profound, pathogen-driven epigenomic reprogramming that is largely unique to the host-adapted virulent pathogen (M. bovis). Crucially, the authors integrate these regulatory networks with existing Holstein-Friesian GWAS datasets to prioritize novel candidate genes (ERBB4, LRCH1, MRTFA, and RNPC3) associated with M. bovis infection susceptibility. This work represents a significant advancement in our understanding of host-pathogen interactions and animal resilience to bovine tuberculosis (bTB).

      Strengths:

      (1) The manuscript addresses a major socioeconomic problem in global livestock agriculture and human zoonotic health. By profiling host-adapted vs. non-host-adapted and live vs. dead bacilli, it provides fundamental insights into mycobacterial virulence mechanisms and evolutionary adaptation.

      (2) The multi-omic approach is robustly executed, and the sample size (n=6 for RNA-seq, n=3 for ChIP/ATAC-seq subgroups) is highly appropriate for primary livestock cell cultures.

      (3) The inclusion of live virulent, live attenuated, non-host-adapted, and killed strains allows the authors to dissect whether host responses are driven by passive PAMP recognition or active, pathogen-directed virulence factors.

      (4) The parallel mapping of four distinct histone marks alongside chromatin accessibility mapping (ATAC-seq) yields a highly refined picture of enhancer and promoter dynamics.

      Weaknesses:

      (1) The profound transcriptional response of the gamma-irradiated (IRR) M. bovis group (3,320 DEGs vs. 2,312 for MTU) is intriguing but lacks a deep biochemical and functional explanation.

      (2) While the paper provides a clear atlas of epigenetic alterations, the underlying mycobacterial effectors driving these specific chromatin alterations remain largely correlative.

    4. Reviewer #3 (Public review):

      Summary:

      Hall et al. use a multi-omics approach to investigate the responses of bovine alveolar macrophages to Mycobacterium tuberculosis infections and the underlying mechanisms that are shared between bacterial family members. The study is of particular importance for multiple reasons, including the impact of bovine infections on food supply (and resulting economic impacts) as well as the use of bovines as a large animal model to investigate the effects of M. tuberculosis infection. The authors isolated bovine alveolar macrophages and exposed them to infection with M. bovis, M. bovis BCG, irradiated M. bovis, or M. tuberculosis. 24 hours post-infection, samples were analyzed by RNAseq, ChIP-seq, and ATAC-seq to look at alterations in the transcriptome and epigenome. Through fluorescence imaging-based analysis, the authors show equivalent infection (bacterial uptake) in all groups, except for the irradiated M. bovis. Principal component analysis of the transcriptomic data demonstrated strong segregation for the M. bovine-infected cells compared to the other groups, which had some intermixing in the PCA. Among the differentially expressed genes, the cytokine IL36G was significantly upregulated across all four groups. This is significant as this cytokine enhanced autophagy in macrophages and subsequent M. tuberculosis killing activity. To further investigate the transcriptional changes, ChIP-seq and ATACseq were utilized to investigate chromatin changes in the form of differential affinity binding sites (DABS) and differential open chromatin regions (DOCR). TPMRSS2, a protease that plays an important role in multiple types of infections (tuberculosis, COVID, etc.), was found to be a significant DOCR in both the M. bovis and M. tuberculosis challenge groups, conferring enhanced TPMRSS2 expression in both of these groups. Using the integrative approach between all the omics data, the authors found that CD274, which encodes the PD-L1 protein, was upregulated in all four groups. PD-L1 is known to play an immunosuppressive role, and PD-L1+ macrophages have been shown to create "cold" microenvironments that would likely favor the mycobacterium. Lastly, SNP analysis found variants in four genes (ERBB4, LRCH1, MRTFA, RNPC3) that could serve as susceptibility genes.

      Strengths:

      Overall, this study demonstrates that challenge with M. bovis elicits a more extensive remodeling of chromatin and subsequent gene expression changes in macrophages compared to the other closely related strains. The work demonstrates the power of functional genomics and its utility in investigating the underlying changes that can affect responses to infections and subsequent outcomes. While the study lacks functional validation, the cohesive dataset is quite compelling, and from it, the authors draw conservative conclusions and are frank about their study limitations.

      Weaknesses:

      Lack of functional validation of some of the targets found.

    1. eLife Assessment

      This important study investigates how verbal and nonverbal working-memory processing is distributed across large-scale functional networks in the human brain using precision fMRI. By leveraging extensive within-subject data and individualized network mapping, the authors provide solid evidence that hemispheric specialization for verbal versus nonverbal information extends across multiple association networks and is reproducible across independent datasets. The use of state-of-the-art precision neuroimaging approaches reveals fine-grained laterality patterns that are likely obscured in conventional group analyses.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Sun et al. investigate the hemispheric lateralization of functional brain networks during verbal versus nonverbal working memory tasks. Utilizing state-of-the-art precision neuroimaging in highly sampled individuals, the authors define a set of distributed association networks and examine their task-evoked responses. The authors report a "generalized laterality effect," wherein multiple association networks appear to functionally split across the hemispheres, with left hemisphere components exhibiting a relative preference for verbal stimuli and right hemisphere components preferring nonverbal stimuli.

      Strength:

      The use of dense-sampling fMRI is a major strength of this study, allowing for a highly accurate, individual-specific mapping of network topologies that group-averaging typically obscures. Despite the interpretational concerns raised below, this high-quality, within-subject imaging dataset represents a valuable resource for the community. Furthermore, the inclusion of an independent prospective replication dataset provides valuable confidence in the robustness of the core imaging metrics. However, while the data quality is exceptionally high, the conceptual interpretations regarding "network splitting", "preferential recruitment", and the generalizability of the verbal/nonverbal dichotomy require significant refinement. Several methodological and statistical clarifications are needed to fully support the authors' claims.

      Weaknesses:

      Major:

      (1) The manuscript relies heavily on task contrast values (e.g., Face > Word) to conclude that networks functionally "split" their profiles, with specific hemispheres being "preferentially recruited" by either verbal or nonverbal materials. While the data clearly demonstrate relative hemispheric differences, claiming absolute bidirectional specialization and active recruitment appears to overstate the findings in two key ways:

      First, statistical evidence for true bidirectional "splitting" is scarce. A significant hemispheric difference confirms a relative shift in processing, but it does not permit claims about absolute preference. When examining the face>word effects against zero in the discovery dataset (Figure 3), the right hemisphere of the LANG, FPN-B, CG-OP, and SAL networks shows no significant preference for faces over words. In the replication dataset (Figure 6), two of the four targeted networks (FPN-A, CG-OP) similarly fail to show a significant right-hemisphere preference. Furthermore, it is unclear whether these tests against zero were corrected for multiple comparisons (e.g., 18 individual tests in Figure 3). Networks reporting significance at the uncorrected p < 0.05 level (such as the left hemisphere of LANG and FPN-B) might not survive standard correction, suggesting that even the left hemisphere's preference for words may be statistically marginal. Second, contrast differences in networks exhibiting negative signals may reflect relative deactivation rather than active recruitment. A mathematically positive contrast value derived from two negative activation states (e.g., Face [-6] > Word [-8]) does not indicate active "recruitment" for face processing. Instead, it merely reflects a relative difference in deactivation. Characterizing this dynamic as "preferential recruitment" misleads the reader regarding the actual physiological state of the network. Furthermore, such asymmetric suppression is frequently driven by generalized differences in task difficulty or cognitive effort, rather than true stimulus-specific processing.

      (2) Building on the previous point, it would be highly beneficial to include the behavioral data (such as accuracy and reaction times) for the N-Back conditions, which do not currently appear to be reported in the manuscript. This information is important because if one condition (e.g., the Face N-back) was significantly more challenging or required greater cognitive effort than the other (e.g., the Word N-back), the observed hemispheric dissociations might reflect differences in arousal, effort, or attentional deployment rather than stimulus-specific processing. ion

      (3) The manuscript claims a "generalized" hemispheric laterality effect. However, the supplemental figures suggest this effect may be highly sensitive to the specific stimuli used in the main text (unfamiliar Faces vs. rhyming Words), which represent extreme ends of visuospatial and phonological processing. As we talked about earlier, a true functional "split" implies that the hemispheres respond in opposite directions. However, visual inspection of the supplemental graphs reveals that for the vast majority of networks, both hemispheres are actually driven in the exact same direction (Figures S8 and S9).

      In the Face > Letter contrast (Figure S8), true bidirectional splits largely disappear. With the exception of FPN-A, the bars for both the left and right hemispheres point in the exact same visual direction for every network (e.g., both hemispheres are visually positive in DN-A and dATN-B, and visually negative in CG-OP and dATN-A). This indicates that the hemispheres actually share the same categorical preference and merely differ in magnitude. Strikingly, the LANG network shows no significant difference between Faces and Letters in either hemisphere, suggesting that the robust leftward shift observed in Figure 3 was possibly driven by the heavy semantic and phonological demands of the rhyming task, rather than a generic preference for "verbal" processing.

      Similarly, in the Scene > Word contrast (Figure S9), the hemispheres do not visually diverge in their response direction for most networks. For example, both hemispheres are visually negative (indicating a shared preference for Words) in the LANG, FPN-B, CG-OP, and SAL networks. Conversely, both hemispheres are visually positive (indicating a shared preference for Scenes) in the dATN-B and DN-A network.

      Because the hemispheres do not visually diverge in their response direction for most networks across these supplementary contrasts, the claim of a robust, generalized hemispheric "split" is unsupported.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Sun and colleagues use precision fMRI to investigate the spatial patterns of verbal versus nonverbal processing in the cortex. They first show that the left hemisphere tends to be more active for word than face processing, whereas the right hemisphere tends to be more active for face than word processing during a working memory task. They then show that this is not confined to a particular network specialized for verbal vs. nonverbal processing but is a pan-network hemispheric difference across 8/9 association networks. This was unexpected, as prior precision imaging work emphasized network specialization and the authors had hypothesized that verbal vs. nonverbal processing would show bilateral network-level effects in a single, right-lateralized network (FPN-B). They then replicated this pan-network effect in an independent dataset.

      Strengths:

      (1) This paper is neat, clear, and succinct. The authors convincingly show that there exists a hemispheric difference in face vs. word processing during a working memory paradigm across many association networks.

      (2) They replicate this result in a fully independent dataset. They do a particularly nice job setting up a prospective replication analysis over 4 distinct association networks.

      (3) They do a wonderful job framing the experiment - they provide context on how previous precision imaging findings would have suggested a specialized network for verbal vs. nonverbal processing, they replicate major findings from that work in this data (network laterality and bilateral network activation in multiple task contexts), and then show the surprising pan-network hemispheric laterality in two independent datasets.

      (4) This paper adds a unique perspective to precision imaging findings. Most precision imaging has shown that different networks seem to be specialized for different cognitive tasks - individual task activations tend to follow network boundaries and networks tend to be cohesively activated. These types of findings led to the idea that previously reported broad hemispheric effects for verbal vs. nonverbal processing could instead be due to a lateralized network specialized for verbal versus nonverbal processing. However, this paper accounts for individualized network topography and yet finds a hemispheric effect that transcends individual networks. I think this result will have a strong influence on how many readers think about network specialization.

      Weaknesses:

      (1) The evidence for the main result comes only from a single fMRI task (n-back) with a particular set of stimuli (faces, words) in which processing demands are not matched across verbal and nonverbal conditions (word blocks use rhyming, face blocks use an exact match). While the claim made is fairly broad (hemispheric laterality in verbal vs. nonverbal processing), it isn't clear how well this finding would generalize to other types of stimuli or processing. This was mentioned in the limitations section, but the paper would be strengthened by both additional evidence for a general verbal versus nonverbal processing effect and by a discussion of how these specific stimuli (words and faces) and processing demands (differences between rhyming and exact match) might affect the results. The authors did include supplemental post-hoc analyses of scene and letter conditions that are matched in processing, but did not discuss them in the main text.

      (2) While lateralization is seen in 8/9 networks, the effect is much larger in some networks (FP-A, DATN-A) versus others. Similarly, the face > word effect is much stronger in some regions of the cortex than others. For example, all the individuals shown exhibit a patch of RH mid-LPFC cortex with particularly strong face > word preference compared to anywhere else in association cortex and an analogous LH mid-LPFC region that shows a particularly strong word > face preference. While the authors are correct that lateralization is seen across many networks, there does seem to be a topography to it rather than a diffuse hemispheric difference. I wonder to what degree the general hemispheric laterality pattern could be driven by a subset of regions that happen to cross several association networks. Neither the differences between networks nor the finer-scale activation patterns are considered in this paper. The paper would be strengthened by considering them.

      (3) The paper is focused on association networks and provides an interesting report on face versus word processing in those networks specifically. In the brain maps, it is possible to see face versus word preference in non-association regions as well. The authors don't discuss these patterns, but the paper would be strengthened by describing the relationship of these patterns with known face and language processing systems in the sensory cortex.

    4. Reviewer #3 (Public review):

      Summary:

      This work takes a precision neuroscience approach to examining how different task demands elicit lateralized vs. bilateral activity in large-scale brain networks. Using ~35-150+ minutes of resting state data per person, the authors identify individualized networks and map their degree of laterality. They find that all association networks are bilateral, with some networks. such as the language network (left) and fronto-parietal B network (right), showing some marked degree of laterality. A sentence reading task evokes activity bilaterally in the language networks, while working memory load during an n-back task evokes bilateral activity largely in fronto-parietal, dorsal attention, and salience networks. Interestingly, though, the authors find that when digging into the N-back task and contrasting blocks that have rhyming words vs. faces, this elicits a more lateralized pattern of activity; left activation for rhyming and right activation for faces. This lateralization is seen across multiple association networks, not just the language or fronto-parietal networks. These findings are then replicated in another precision data set.

      Strengths:

      This work has several notable strengths. This study boasts a lot of data for each individual, allowing them to examine individualized functional networks and task activations. Given the marked individual differences in laterality (see Figures S4 and S5), a group-averaged network atlas and group-averaged activation maps would likely muddle some of the interesting effects.

      The authors also use an elegant approach for individualized network estimation, multisession hierarchical Bayesian modeling, or MSHBM. MSHBM allows vertex-based functional connectivity to steer individualized networks while also incorporating meaningful priors to identify comparable networks across individuals. This is a nice approach combining elements of a common atlas across individuals with data-driven approaches.

      The authors have a rich working memory N-back task with four different stimuli types with which they can examine processing different types of stimuli.

      A major strength of this work is the replication. The authors take their surprising findings and replicate the entire experiment in another sample. The replication sample notably has even more data per person than the discovery sample, allowing for robust and precise estimation of individual networks and task activity.

      Weaknesses:

      I'd like to frame this section as 'unsolved challenges' rather than 'weaknesses'.

      One of the strengths of this paper is that the working memory task that incorporates so many different stimuli and conditions also poses a caveat for interpreting the task stimuli effects. The 'word' condition requires multiple cognitive demands - working memory and rhyming/phonological processing. This makes it a little hard to draw complete parallels between the word and face conditions. It does, however, provide a useful backdrop to explain why there are such strong left laterality effects for the word condition. I would expect that explicit phonological processing that is required for a rhyming task would elicit more left lateralized activation than the sentence processing task, as fluent readers are likely not using explicit phonological processing for sentence processing. This could explain why, for the sentence reading task, they find more bilateral activation, but a task that specifically targets phonological processing (i.e., the 2-back word condition) would evoke more left lateralized activation above and beyond the activation supporting working memory, which is held constant across both the 'words' and 'faces' conditions. Ultimately, I believe this supports the broad idea of this work - that large-scale association networks can show lateralization, even beyond language network boundaries. However, the fact that there is a dual-task load that evokes phonological processing for the word condition is important for contextualizing these findings.

      When looking at plots of individual differences, it is clear that some people are more bilateral than others in certain networks (and sometimes overall). Because of the amount of data per person in this precision study, it naturally limits the overall sample size (N=29). This makes it difficult to interrogate individual differences in laterality and what might predict this effect across people.

    1. eLife Assessment

      Murrell et al. describe a high-throughput method, the Feeding Experimentation Device (FED3), to study food foraging strategies in mice. The authors provide solid evidence about key sex differences in foraging strategies, specifically the finding of greater male win-stay behavior. Further consideration of how sex differences could be uniquely influenced by FED3 testing conditions (e.g., single-housing, hormones) and task demands (e.g., 100-0 vs. 80-20) would be helpful. Given this open source FED3 platform, the authors provide valuable findings that have utility to the field of behavioral neuroscience, specifically to those interested in studying reward-related behavior.

    2. Reviewer #1 (Public review):

      Summary:

      This paper examines potential sex differences in the conflict between exploitation, pursuing food and rewards in previously-associated locations/paradigms, or exploration of new locations that might result in better outcomes. Dysregulation of this conflict may be an underlying behavioral modality of psychiatric diseases. They used four distinct tasks: a two-armed Bandit 100:0 task, a standard fixed ratio 1 task, a two-armed Bandit 80:20 task, and a closed-loop economy PR1 task that allows for the assessment of motivational breakpoint.

      Male mice show significantly higher accuracy under conditions of high probability known rewards, sticking with an action that just resulted in a reward or "win-staying". This was demonstrated in multiple paradigms, and there was a predictive nature of this behavior that could predict animal sex with modest accuracy. Under probabilistic environments, males were no longer more accurate than females but still used a higher win-stay strategy. A closed-loop PR1 task showed that there were no inherent differences in motivational breakpoint between sexes. Finally, the authors use simulations to determine an appropriate number of animals needed to detect these differences.

      Strengths:

      The manuscript attempts to resolve inconclusive sex differences that have heretofore been neglected or inconclusive due to insufficient power. The most impressive aspect of this paper is its scale, assaying 62 female mice and 74 male mice in identical exploration-exploitation tasks using high-throughput and noninvasive operant feeding via FED3. Very few labs can achieve this scale, which is necessary to detect sex differences with a small effect size.

      The authors use some sophisticated modeling approaches and analysis of data from the 136 mice to investigate the significance of these sex differences and interrogate other conditions. They also use simulations to model the likelihood of replicating these differences given a sample size. This is extremely helpful for other researchers as they consider sex as a biological variable.

      Weaknesses:

      The study is largely descriptive in nature and does not pursue any mechanism of the underlying differences, like hormones, neuromodulators, or circuits. The lack of estrous cycle tracking is acknowledged as a limitation.

    3. Reviewer #2 (Public review):

      Summary:

      Murrell and colleagues examine sex differences in mouse decision-making tasks, using the FED3 device, which allows for continuous data collection in the home-cage. Mice performed four tasks across two weeks, which provided all of their food. Across tasks, male mice were more likely to repeat a rewarded choice than females, which benefits decision accuracy in deterministic tasks. This work complements existing results for decision-making differences in males and females, affirming that this domain of cognition is particularly sensitive to sex differences. However, there are some specific features of the FED3 device, such as single housing, closed economy feeding, and 24-hour access that can uniquely influence decision-making in a (likely) sex-dependent manner that encourage considering these data as examining sex differences in a particular context, rather than as a generalized finding. At the same time, these data could offer new insights about nuances of behavior like circadian rhythms or bout analysis, uniquely enabled by the extended availability of the FED3 devices. The analyses in this paper also make an important point, encouraging researchers to use methods that allow for much larger N's to provide clearer and more robust results.

      The FED3 devices are an innovative new way to approach behavior, and have allowed the authors to test many dozens of mice in a battery of tasks, over which they see similar patterns of increased win-stay behavior in male C57b6 mice (wildtypes from several knockout lines). The authors point out that there are discrepancies in prior literature across tasks and species in terms of how sex differences influence decision making, but there are some particular ways that sex differences could interact with the FED3 devices that it would be interesting and important to consider further. In particular, the fact that the animals live with the device, singly housed, may be an underrecognized contributor to sex differences. Changes in social interaction and dominance arising from long-term single housing are very likely to impact males and females differently, for example.

      Continuous data collection is a fascinating way to look at learning and decision-making, but it also raises interesting questions about whether these dynamics are impacted as a function of continuous access to the device. In addition to summary metrics over the whole task, it might be valuable to look at learning across the task each day, and within circadian periods of each day. For example, it seems based on the example sessions for the 100-0 bandit task that animals take at least a few reversals to learn the task structure. How many trials does it take for a male or female mouse to reach some criteria of success? Do the sex differences exist at all time points? Does the light cycle affect the accuracy or trial counts? There are numerous such analyses that could particularly inform future use of the FEDs across laboratories, and identification of similar or distinct patterns of sex differences or behavior in other apparatuses, and would be a benefit to the field.

      The authors employed several computational techniques to identify parameters or features of behavior that might explain the sex differences they observed, and this is a strength of the manuscript. However, the win-stay lose-shift agent may not be an ideal match to make conclusions about exploitation, as it is unclear how win-stay and lose-shift strategies map onto explore/exploit tradeoffs. If an animal were exploiting an option, they may win-stay *and* lose-stay, if the task is probabilistic. Indeed, the model fit is weaker for the 80-20 bandit, suggesting this model may not reflect the actual strategies mice are engaging in, potentially in both bandit tasks. This point is particularly worth considering in light of the criteria for shifts in the tasks being based not on the number of trials completed, but on the number of pellets earned. When shifts are tied to reward collection rather than trials, it can amplify differences in behavior driven by reward consumption.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Murrell et al. describes a high-throughput approach for evaluating food foraging strategies in mice. Building on their prior publication describing the technical aspects of the Feeding Experimentation Device 3 (FED3), this study demonstrates the utility of the FED3 in evaluating decision-making in mice. The authors identify key differences in male and female foraging strategies that could not be accounted for by total food consumption or overall food motivation. Given the rapid adoption of the open source FED3 platform, this work is likely to be of broad interest and utility in the field.

      Strengths:

      (1) The use of cost-effective, open-source devices like FED3 provides substantial value to the scientific community. Validation of appropriate conditions for using this equipment is an important step toward broad adoption.

      (2) The authors implement a simple but elegant experimental design for studying food-motivated decision-making behavior. This approach could be applied to a wide range of preclinical disease models in future studies.

      (3) The study is well-powered to evaluate sex as the primary experimental factor (62 females, 74 males), allowing the authors to make convincing claims about differences in strategy. Additionally, the dataset provides a useful benchmark for power analyses in future studies involving more complex experimental designs.

      (4) The figures are clear and generally easy to interpret the primary findings.

      (5) The conclusions are appropriate and not overstated

      Weaknesses:

      (1) A major strength of this study is the potential utility for new investigators trying to implement cognitive behavioral tasks in mice. However, the present version provides limited background on the rationale for selecting the bandit task and on prior work applying it in similar contexts. Including additional background and discussion would better contextualize the approach for other groups considering adopting it in their own studies.

      (2) Some methodological details surrounding the initiation of the experiment could be clarified. Specifically, it is unclear if mice transitioned directly from standard housing conditions (group housed, standard chow) to the study conditions (single housed, FED3-based probabilistic learning), or intermediate acclimation/training steps were used, such as autoshaping, free access to new food pellets, or FR1 training. A more detailed experimental timeline (for example, see Figure 1 from PMID 39710132) would address this concern.

      (3) The authors evaluated multiple probabilistic conditions (100%, 90%, 80%, 70%, 60%), but ultimately focused on the 80% condition for this study. A more detailed explanation for how this conclusion was reached would be useful for future researchers working under different experimental conditions (i.e., age, strain, genotype, disease model) where other probabilistic conditions may be more appropriate.

    1. eLife Assessment

      This study sets out to show that the directionality of bacterial transport in confined environments is affected by a salt gradient. Using Pseudomonas putida in microfluidics devices, the authors hypothesize that salt-induced diffusiophoresis acts as a steering mechanism to guide microbial movement. However, the evidence to support this hypothesis is incomplete, because the underlying reorientation process is not directly demonstrated, and the necessary controls to rule out alternative hypotheses are not presented.

    2. Reviewer #1 (Public review):

      Summary:

      The authors claim that bacteria are guided by diffusiophoresis. They perform experiments of bacterial motility in microfluidic channels with salt gradients. The data show that P. pudita bacteria swim towards higher sodium chloride concentrations, but there is no evidence that this is due to a diffusiophoresis.

      Weaknesses:

      It is well known that bacteria perform chemotaxis in salt gradients (see e.g., PNAS 86, pp. 8358-8362, 1989). The underlying mechanism based on chemoreceptors is widely accepted, but the authors do not mention this possibility. I recommend a control experiment where the chemotaxis genes are knocked out. Even if this mechanism can be ruled out, the current data show no evidence for a mechanism based on diffusiophoresis.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigate how salt gradients influence the transport of Pseudomonas putida in confined microfluidic environments. They report that salt gradients enhance directional migration, increase run persistence, and promote transport toward contaminant-rich regions. To explain these observations, the authors propose a physical steering mechanism in which differential diffusiophoretic mobilities of the cell body and flagellar bundle generate an aligning torque that reorients cells along the salt gradient.

      Strengths:

      The study addresses an interesting question at the interface of microbiology, complex fluids, and active matter. Their experiments suggest that salt gradients influence bacterial transport behavior and lead to more persistent, directional motion. Once confirmed, the proposed mechanism would broaden our understanding of how environmental gradients can shape microbial migration through physical interactions in addition to more traditional sensing-based pathways.

      Weaknesses:

      The main limitation of the current study is that the proposed steering mechanism is not directly demonstrated. The evidence for the diffusiophoretic torque is largely inferred from trajectory statistics and theoretical modeling. While the observed transport behavior is convincing, the causal link between the observed migration patterns and the proposed reorientation mechanism remains less well established. In particular, the manuscript focuses primarily on cell trajectories and transport properties, whereas the proposed mechanism fundamentally involves changes in cell orientation. Additional evidence connecting orientation dynamics to the proposed torque mechanism would strengthen the conclusions.

      A related concern is whether alternative physical mechanisms associated with the imposed salt gradients have been fully excluded. For example, weak flow-mediated effects or other hydrodynamic influences could potentially contribute to the observed transport behavior. The manuscript would benefit from a more thorough discussion of such possibilities and a clearer justification for why the proposed diffusiophoretic mechanism should be regarded as the dominant explanation.

      The manuscript would also benefit from a clearer positioning within the broader literature on physically induced microbial transport and swimmer reorientation. Previous studies have demonstrated directed migration arising from rheotaxis (Marcos et al., 2012, PNAS) and viscosity-gradient-induced steering (Stehnach et al., 2021, Nature Physics). While the mechanism proposed here appears distinct, a more explicit discussion of how the present work relates to these earlier studies would help readers better understand the specific conceptual advance being made.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors claim that bacteria are guided by diffusiophoresis. They perform experiments of bacterial motility in microfluidic channels with salt gradients. The data show that P. putida bacteria swim towards higher sodium chloride concentrations, but there is no evidence that this is due to diffusiophoresis.

      Weaknesses:

      It is well known that bacteria perform chemotaxis in salt gradients (see e.g., PNAS 86, pp. 8358-8362, 1989). The underlying mechanism based on chemoreceptors is widely accepted, but the authors do not mention this possibility. I recommend a control experiment where the chemotaxis genes are knocked out. Even if this mechanism can be ruled out, the current data show no evidence for a mechanism based on diffusiophoresis.

      We thank the reviewer for raising this important comment. We agree that receptor-mediated salt taxis is a well-established mechanism in some bacteria, including the classic study by Qi and Adler (PNAS, 1989). We note, however, that the original manuscript did discuss this possibility and cited Qi and Adler in line 117: “We also note that we did not observe any significant difference in the tumble rates between the control and NaCl gradient cases (Figure 3h; Figure S2, SM), suggesting that NaCl gradients do not interfere with chemoreceptors (Qi and Adler, 1989).” In Figure S2, the run-time distributions show no significant difference between the no-gradient condition and the NaCl-gradient condition, with fitted tumble rates of λ = 0.41 s<sup>−1</sup> and λ = 0.38 s<sup>−1</sup>, respectively. We also do not observe a directional bias in run duration, namely longer runs up the salt gradient and shorter runs down the gradient, which would be expected for canonical chemoreceptor-mediated taxis.

      The basis for assigning the observed migration to diffusiophoresis is that the NaCl gradient produces a directional drift of the bacterial body without a measurable change in the run-and-tumble statistics. This behavior is consistent with our previous work [1], where non-motile bacteria were shown to undergo diffusiophoretic migration toward higher salt concentration. Because that migration occurred in non-motile cells and across different bacterial types and morphologies, it supports the interpretation that native bacterial surface charge can drive a non-specific diffusiophoretic response in salt gradients.

      That said, we agree with the reviewer that a genetic control would provide a stronger test against receptor-mediated chemotaxis. We will therefore perform additional experiments using a ∆cheA strain. Because CheA is required for canonical chemotactic signal transduction, observing the same directional migration in the ∆cheA mutant would directly test whether the NaCl gradient response persists in the absence of receptor-mediated chemotaxis. We will include these new data and revise the manuscript to more explicitly distinguish diffusiophoretic drift from chemoreceptormediated salt taxis.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate how salt gradients influence the transport of Pseudomonas putida in confined microfluidic environments. They report that salt gradients enhance directional migration, increase run persistence, and promote transport toward contaminant-rich regions. To explain these observations, the authors propose a physical steering mechanism in which differential diffusiophoretic mobilities of the cell body and flagellar bundle generate an aligning torque that reorients cells along the salt gradient.

      Strengths:

      The study addresses an interesting question at the interface of microbiology, complex fluids, and active matter. Their experiments suggest that salt gradients influence bacterial transport behavior and lead to more persistent, directional motion. Once confirmed, the proposed mechanism would broaden our understanding of how environmental gradients can shape microbial migration through physical interactions in addition to more traditional sensing-based pathways.

      Weaknesses:

      The main limitation of the current study is that the proposed steering mechanism is not directly demonstrated. The evidence for the diffusiophoretic torque is largely inferred from trajectory statistics and theoretical modeling. While the observed transport behavior is convincing, the causal link between the observed migration patterns and the proposed reorientation mechanism remains less well established. In particular, the manuscript focuses primarily on cell trajectories and transport properties, whereas the proposed mechanism fundamentally involves changes in cell orientation. Additional evidence connecting orientation dynamics to the proposed torque mechanism would strengthen the conclusions.

      We thank the reviewer for the constructive comments. We agree that the proposed steering mechanism should be supported by evidence that directly connects the salt gradient response to bacterial orientation dynamics, not only to trajectory-level transport statistics.

      We would like to clarify that the original manuscript already includes an orientation-based analysis in Figure 4f,g in the main text. The corresponding methodology and results are described in lines 173–183 and in the Supporting Information. Specifically, we quantified cell steering by measuring the change in body angle, ∆θ, along individual run trajectories as a function of arc length, s, using the orientation correlation ⟨cos(∆θ)⟩<sub>s</sub>. In the absence of salt gradients, the orientation correlation decays slowly with arc length, indicating persistent swimming along the initial

      Author response image 1.

      Instantaneous angular velocity as a function of heading angle relative to the salt gradient orientation. (a) Experimental and (b) simulated mean angular velocity of cells as a function of heading angle θ (measured relative to the gradient direction; θ = 0° points toward the gel/high-salt side) in the absence (blue) and presence (red) of a NaCl gradient. Positive and negative values indicate counterclockwise and clockwise rotation, respectively, with arrows showing rotation direction. Under the gradient, both experiments and simulations show a signed, angle-dependent rotation rate that is largest near θ = ±90° and approaches zero near θ = 0° and 180°, consistent with a restoring torque that steers cell heading toward the gradient direction. Simulations reproduce this behavior with comparable magnitude to the experimental measurements, and in the absence of a gradient, angular velocity remains relatively small in the no-gradient case with no consistent directional bias across heading angles.

      run direction. Under a salt gradient, the correlation decays more rapidly, indicating stronger directional reorientation during runs. Because the tumble statistics do not change significantly between the no-gradient and salt-gradient conditions, this enhanced orientational decorrelation is not attributed to increased tumbling or rotational noise. Instead, it is consistent with continuous deterministic steering during runs, as expected from a diffusiophoretic torque acting on the cell body–flagellar bundle system.

      To further address the reviewer’s concern, we performed an additional orientation-dynamics analysis using the same dataset shown in Figure 4. Following the approach used by Stehnach et al. [2], we calculated the instantaneous angular velocity during individual runs as a function of the cell heading angle relative to the salt gradient direction. The cell orientation was obtained from the run trajectories, and tumble events were excluded because they produce large transient angular velocity spikes that are not representative of continuous steering during runs.

      The new analysis is shown in Author response image 1. Under the no-gradient condition, the angular velocity remains small and nearly independent of heading angle. In contrast, under the salt gradient condition, the angular velocity becomes strongly heading-dependent. The angular velocity is largest when cells swim nearly perpendicular to the salt gradient, where a steering torque is expected to be maximal. Moreover, the sign of the angular velocity indicates rotation toward alignment with the gradient direction. This behavior is consistent with the proposed diffusiophoretic torque mechanism and provides a direct link between the observed transport behavior and salt-gradient-induced reorientation dynamics.

      We plan to add this angular velocity analysis to the revised manuscript and revise the relevant text to make the connection between trajectory statistics, orientation dynamics, and the proposed torque mechanism clearer.

      A related concern is whether alternative physical mechanisms associated with the imposed salt gradients have been fully excluded. For example, weak flow-mediated effects or other hydrodynamic influences could potentially contribute to the observed transport behavior. The manuscript would benefit from a more thorough discussion of such possibilities and a clearer justification for why the proposed diffusiophoretic mechanism should be regarded as the dominant explanation.

      We thank the reviewer for raising this important point. We agree that alternative physical mechanisms associated with the imposed salt gradient should be considered explicitly. In the revised manuscript, we will add a more detailed discussion explaining why flow-mediated or other hydrodynamic mechanisms are unlikely to account for the observed steering behavior.

      First, the characteristic diffusiophoretic drift velocity in our experiments is approximately u<sub>d</sub> ≈ 1 µm/s, corresponding to only about 1–5% of the typical swimming speed of P. putida. Thus, the proposed mechanism does not require externally driven advection of the cells. Instead, the salt gradient produces a weak but persistent differential diffusiophoretic slip on the cell body and flagellar bundle, which can generate a reorienting torque during active swimming.

      We also considered whether diffusio-osmotic flow along the channel walls could generate sufficient shear to induce rheotaxis. In a dead-end channel, the diffusio-osmotic velocity profile can be estimated as [3]

      which gives a wall shear rate (at z = h) of

      This value is below the shear rate threshold reported by Marcos et al. [2], where rheotactic drift becomes negligible for S < 0.1 s<sup>−1</sup>. Therefore, the shear generated by diffusio-osmotic wall flow in our experiments is too weak to explain the observed directional reorientation. This effect would be even smaller for P. putida, whose thin flagellar bundle is expected to experience weaker shear-induced alignment than organisms with larger flagellar structures.

      We further considered viscotaxis as a possible mechanism. However, the viscosity difference between 1 mM and 100 mM NaCl solutions is marginal. This contrast is far smaller than the approximately 4–5-fold viscosity difference reported to induce strong viscophobic turning in bacteria [4]. Thus, salt-gradient-induced viscosity variations are insufficient to account for the measured steering response.

      Taken together, these estimates indicate that hydrodynamic shear, rheotaxis, and viscotaxis are too weak under our experimental conditions to explain the observed migration and orientation dynamics. In contrast, the proposed nonuniform diffusiophoretic mechanism naturally accounts for the key observations: directional migration up the salt gradient, enhanced curvature during runs, heading-dependent angular velocity, and the absence of significant changes in tumble statistics. We will include this analysis in the revised manuscript to clarify why diffusiophoretic steering is the dominant mechanism under the present conditions.

      The manuscript would also benefit from a clearer positioning within the broader literature on physically induced microbial transport and swimmer reorientation. Previous studies have demonstrated directed migration arising from rheotaxis (Marcos et al., 2012, PNAS) and viscosity-gradient-induced steering (Stehnach et al., 2021, Nature Physics). While the mechanism proposed here appears distinct, a more explicit discussion of how the present work relates to these earlier studies would help readers better understand the specific conceptual advance being made.

      We thank the reviewer for this helpful suggestion. We agree that the manuscript should more clearly position the proposed mechanism within the broader literature on physically induced microbial transport and swimmer reorientation.

      In the revised manuscript, we have added a discussion comparing our results with prior studies on rheotaxis and viscosity-gradient-induced steering. Specifically, we now discuss the work of Marcos et al. [4], which showed that shear flow can generate a torque on the helical flagellum of Bacillus subtilis, producing rheotactic alignment independent of chemical sensing. We also discuss the work of Stehnach et al. [2], which showed that viscosity gradients can steer Chlamydomonas reinhardtii through asymmetric viscous drag on its two flagella, producing viscophobic turning down the viscosity gradient.

      The mechanism proposed in the present work is distinct from both of these cases. Unlike rheotaxis, it does not require externally imposed shear flow. Unlike viscophobic turning, it does not rely on a substantial viscosity contrast. Instead, we propose that a salt concentration gradient generates differential diffusiophoretic motion of the cell body and flagellar bundle, producing a torque that continuously reorients swimming cells along the gradient. This mechanism therefore identifies salt gradients as a distinct physical cue capable of steering bacteria through surface-mediated transport rather than through flow, viscosity contrast, or canonical chemoreceptor signaling.

      We have added the following text to the revised manuscript: ”Our findings also relate to other physical mechanisms of microbial reorientation. Bacterial rheotaxis, arising from a torque generated by shear flow acting on the helical flagellum, steers Bacillus subtilis independent of any chemical gradient [4]. Viscosity gradients similarly drive a viscophobic turning in the alga Chlamydomonas reinhardtii, where uneven viscous drag on its two flagella produces a torque that reorients cells down the gradient [2], a behavior confirmed by measuring angular velocity as a function of heading angle and revealing a sinusoidal form, ω(θ) = −ω <sub>visc</sub> sin(θ), which resembles what we report here for diffusiophoresis in (Figure S4). This identifies salt gradients, independent of flow or viscosity, as a distinct physical route by which swimming cells can be steered and guided, broadening the set of known nonchemoreceptor mechanisms for directed microbial transport.” References

      (1) V. S. Doan, P. Saingam, T. Yan, S. Shin, A trace amount of surfactants enables diffusiophoretic swimming of bacteria, ACS Nano 14 (10) (2020) 14219–14227.

      (2) M. R. Stehnach, N. Waisbord, D. M. Walkama, J. S. Guasto, Viscophobic turning dictates microalgae transport in viscosity gradients, Nature Physics 17 (8) (2021) 926–930.

      (3) S. Shin, E. Um, B. Sabass, J. T. Ault, M. Rahimi, P. B. Warren, H. A. Stone, Size-dependent control of colloid transport via solute gradients in dead-end channels, Proc. Natl. Acad. Sci. 113 (2) (2016) 257–261.

      (4) Marcos, H. C. Fu, T. R. Powers, R. Stocker, Bacterial rheotaxis, Proceedings of the National Academy of Sciences 109 (13) (2012) 4780–4785.

    1. eLife Assessment

      This important manuscript presents a way to understand allostery and allostery-inspired mechanical systems. It presents a mechanochemical model of an ATPase-like molecular machine in which geometric constraints and mechanically coupled linkages generate negative allostery, yielding a large stochastic reaction network that captures productive and futile cycling behavior. The work presents an interesting framework for bridging concepts from molecular motors and allostery. It provides a compelling, intuitive physical mechanism for mechanochemical transduction and reveals nontrivial tradeoffs that emerge from geometry-induced constraints.

    2. Reviewer #1 (Public review):

      Summary:

      This paper presents a creative physics-based model of an ATPase-like molecular machine using mechanically coupled linkages to mimic allosteric cycles. The authors construct the complete reaction network by enumerating all mechanically allowed binding configurations and transitions between them. The resulting system contains hundreds of microstates connected through node-level binding, dissociation, intramolecular rearrangements, cleavage, and ligation reactions. Stochastic simulations are then used to study how the machine cycles between ligand-bound and substrate-bound states. Overall, the manuscript presents an interesting and creative mechanochemical framework for modeling ATPase-like allosteric cycles integrating multivalent binding, geometric exclusion through rigidity arising from binding of substrates and ligands, with stochastic simulations.

      Strengths:

      (1) The manuscript presents a creative mechanochemical framework that combines multivalent binding, geometric exclusion, rigidity-based coupling, stochastic kinetics, and catalysis within a unified model.

      (2) The use of geometric exclusion and rigidity to generate negative allosteric coupling is elegant and provides an intuitive physical mechanism for coordinated molecular behavior.

      (3) The interpretation of catalysis as a transient release of mechanical constraints is conceptually interesting and offers a novel perspective on how energy-consuming reactions can regulate state transitions.

      (4) The distinction between productive and futile cycles is insightful and provides a useful framework for understanding pathway selection in molecular machines.

      (5) The explicit construction of a stochastic state network allows the authors to connect microscopic binding events with emergent cyclic behavior.

      (6) The work provides a conceptual platform for exploring how simple mechanical principles may give rise to allosteric regulation and mechanochemical transduction in synthetic molecular systems.

      Weaknesses:

      (1) The manuscript is very dense and difficult to follow. The notation and microstate labels (e.g., {S/L}:{10,5}) obscure the central ideas, and the stochastic model is not explained clearly enough. The authors should provide a simpler schematic, a mapping of state labels, and a step-by-step example of a productive cycle. The supplementary videos would also benefit from additional explanation.

      (2) The framework appears most applicable to mechanically gated motor proteins and may not generalize to allosteric enzymes that operate through conformational ensembles, dynamic coupling, or entropy-driven regulation. The scope of the model should be discussed more carefully.

      (3) The reported behavior appears highly dependent on specific parameter choices and rate hierarchies. A broader sensitivity analysis is needed to demonstrate robustness.

      (4) The binary treatment of states as either rigid or flexible oversimplifies the continuous energy landscapes and fluctuations observed in real biomolecules. The limitations of this approximation should be discussed.

      (5) The role of nonequilibrium thermodynamics is underdeveloped. The relationship between the model, ATP chemical potential, free-energy dissipation, entropy production, and the energetic cost of futile cycles should be discussed more explicitly.

      (6) Although a large number of microstates are enumerated, it remains unclear which states and pathways dominate the dynamics. A more coarse-grained analysis highlighting the key states and transitions would improve interpretability and facilitate comparison with experimental systems.

    3. Reviewer #2 (Public review):

      This is an interesting study aiming to capture the fundamental principles of ATPase-like machines with an elementary model of rigid bars. While molecular motors have been the subject of many studies in statistical physics, taking very simplified approaches, these past studies generally abstract away from geometrical constraints and do not account for allosteric mechanisms. In turn, several simple physical models of allostery are now available, but most consider only long-range effects without reference to reactivity or the conversion of chemical to mechanical energy. The essential ingredient of the present model is a form of negative allostery that stems from geometrical constraints. These constraints impose the presence of hidden microstates and the connections between states that form a reaction network. Nontrivial tradeoffs can then be derived, e.g., on the catalytic rate.

      One limitation of the approach is that energetic constraints, which are equally important, are not themselves derived from physical principles. This includes the different kinetic rates, the binding constants, and the mechanisms of reactivity, even though, in principle, they should follow from basic interaction energies in the context of thermal fluctuations. This is a legitimate choice of modeling level, handled notably by imposing a hierarchy between rates. It would be appreciated, however, if the overall logic of the derivation were clarified, presenting more clearly from the beginning what the fundamental hypotheses of the model are, which aspects are derived from these hypotheses, and which require additional assumptions. It seems indeed that the model involves both fundamental physical assumptions that are used to derive some emerging kinetic features (e.g., geometry imposes futile cycles) and global kinetic assumptions that are used to derive some microscopic features (e.g., a productive cycle imposes the relative values of the kinetic rates).

      The conclusion ends with a proposal to generalize to more elaborate models. I was wondering, however, if the opposite would not be desirable. The current model is already quite involved (as the 7 pages of SI listing transitions between states testify). Wouldn't it be possible to obtain some of the main results, e.g., the trade-offs defining an optimal cleavage rate, from an even simpler model?

    4. Reviewer #3 (Public review):

      Summary:

      In "Design of a minimal, allosteric, and ATPase-like machine using mechanical linkage," Omabegho introduces a simplified model of an allosterically-regulated molecular machine. The model machine takes inspiration from simple ATPase motors, specifically myosin or dynein, and is meant to capture the interactions between enzyme, ligand, and substrate that empower these enzymatic biomolecular machines. The model system attempts to specifically model the inhibitory allosteric regulation relevant to machine operation, following work on mechanical features of allostery in protein-analogue elastic network models. After the introduction of the model machine, the manuscript discusses the primary cycle by which the substrate serves to displace and then cyclically replace the ligand (roughly corresponding to a "recovery" and "power" stroke, respectively, in their biomolecular machine inspiration), as well as cataloguing other futile cycles or individual chemical pathways the machine may traverse. After an exhaustive enumeration of all possible states and transitions of the model machine, the actual mechano-chemical system is mapped to an abstracted stochastic chemical reaction network, which is finally studied numerically in more detail. We believe this model system is novel, interesting, and, importantly, interpretable and could prove to be useful in future modeling and rational design of simple bio-mimetic nanoscopic systems.

      Strengths:

      (1) Omabegho introduces a novel chemo-mechanical model system that captures the core behavior of a simple ATPase molecular machine. The model is relatively simple in construction, but, as the manuscript demonstrates, it can display very rich behavior depending on various competing chemical timescales and mechanisms. These features suggest that this system could, in fact, serve as a useful starting model if adopted by the wider community.

      (2) A key feature to highlight is the mechanistic interpretability of the model. By cutting to the seemingly core functional details of an ATPase-like machine, Omebegho is able to algorithmically produce an exhaustive listing of all allowed chemical states and all possible transitions between them, enabling the study of all possible mechanistic pathways and the relative frequency of each observed path. Further, Omebegho produced very clear visualizations of these states, transitions, pathways, and cycles that again facilitate the interpretability and utility of the model. This interpretability is key to our suggestion that this model could be more widely adopted as a clear ATPase analogue model for study by the broader biophysical community.

      Weaknesses:

      There are some issues to note.

      (1) First, although the model is inspired by ATPases, the cycle in question is not a model even for the significantly abstracted form presented in Supplemental Section 1.1, as noted in the manuscript. The allosteric regulation utilized in the model does not mandate that the ligand (a stand-in for actin, for example) displaces both products (stand-ins for ADP and P). Rather, one product (P) dissociates spontaneously, and the other, larger one (ADP) is allosterically displaced. We recognize mandating this small further detail would presumably complexify the model (i.e., it may require an additional tile in the enzyme construction), but it would also lead to a more accurate, while still simple and interpretable, picture of the biological molecular machines in question.

      (2) Much as we recognize and applaud the model's structure as simple and interpretable, when trying to actually study the chemical dynamics and observed dynamical behaviors, individual chemical rates of various binding and unbinding processes appear to be quite finely tuned. There is some attempt at studying the behavior of this model in various parameter tuning regimes, but the large number of model parameters makes a truly complete numerical study prohibitive. Given this difficulty, it would be instructive to know if a comparison could be made to the actual motivating biological systems and any observed chemical rates in the experiment. Further, given the desired cyclic behavior, it would be quite interesting to see to what degree this behavior is robust to various parameter choices. Again, preliminary work was done in the manuscript by biasing rates or numerical siloing experiments, but a far more exhaustive study is certainly worthwhile.

      (3) Perhaps our biggest critique of the manuscript is that, although it incorporates both mechanical and chemical aspects into the model system construction, all mechanical aspects of the model simply function to limit allowed state transitions. From our understanding, all mechanical aspects of the model are, in fact, abstracted away during any simulations. This modeling choice, of course, retains some mechanical inspiration while making the resulting system more tractable, but it is ultimately not an actual mechanical model. ATPases, especially motors such as myosin and dynein, are inherently mechanical, their fundamental features being the conversion of chemical fuel into mechanical motion. A fuller treatment, perhaps in future work, should include these physical degrees of freedom in simulation, thus truly tracking the interplay between physical enzyme mechanics and chemistry. The absence of actual mechanics as yet, beyond a strict allosteric restriction, is an inherent limitation of the model.

    1. eLife Assessment

      This study presents a useful application of trans-omic network analyses to existing human brain datasets, generating systems-level insights into metabolic dysregulation in Alzheimer's disease. Overall, the authors' analytical choices are solid, with appropriate use of existing data and methods, and with many of their results confirming previous findings. However, some of the authors' key claims, related to previously unknown details of regulatory relationships, are only partially supported due to limitations in dataset cell-type resolution and network robustness, as well as a lack of functional validation. This work will be of interest to cellular or systems neurobiologists studying Alzheimer's disease and could serve as a helpful starting point for future work.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      The authors aim to reconstruct a multi-layer metabolic regulatory network in Alzheimer's disease by integrating transcriptomic, proteomic, and metabolomic datasets from human brain tissue. By linking transcription factors, enzyme expression, metabolic reactions, and metabolites, the study seeks to provide a systems-level understanding of disease-associated metabolic dysregulation.

      The revised manuscript has improved in clarity and includes additional analyses in response to prior comments, particularly the incorporation of cell-type proportion estimates derived from matched single-cell datasets. This represents a meaningful step toward addressing concerns about the interpretation of bulk tissue and strengthens the study's descriptive rigor.

      However, several key limitations remain. First, although cell-type proportions are now estimated, this information is not integrated into downstream analyses. As a result, it remains unclear whether the observed metabolic changes reflect cell-intrinsic regulation or shifts in cellular composition. This distinction is critical for interpreting the inferred regulatory network and limits the strength of the conclusions.

      Second, the biological conclusions remain largely confirmatory. The reported downregulation of energy-related pathways, including the TCA cycle and oxidative phosphorylation, is consistent with prior literature. The trans-omic framework provides a structured representation of these changes, but the manuscript does not convincingly demonstrate that this approach yields new mechanistic insight beyond existing knowledge. In particular, the network is not sufficiently leveraged to identify novel regulatory relationships or generate testable hypotheses.

      Third, while the framework integrates multiple molecular layers, the contribution of upstream transcriptional regulation remains unclear. The most compelling findings appear to arise from protein and metabolite layers, and the manuscript does not clearly demonstrate how transcription factors or mRNA-level changes contribute to the interpretation of metabolic dysregulation. This weakens the claim that the study provides a fully integrated trans-omic perspective.

      Fourth, concerns regarding network robustness remain. The analysis relies on partially overlapping cohorts across omics modalities, and although this limitation is acknowledged, no formal sensitivity or robustness analyses are presented. This reduces confidence in the stability and generalizability of the inferred network structure.

      Finally, the handling of covariates remains limited. While the authors justify this based on data availability and prior studies, the lack of consistent adjustment for known confounders introduces uncertainty in attributing observed differences specifically to disease-related biology.

      Overall, the study presents a technically sound application of a trans-omic integration framework and provides a coherent overview of metabolic dysregulation in Alzheimer's disease. However, the findings are primarily descriptive and confirmatory, and the current analyses do not fully demonstrate that the approach yields novel biological insight. The work will be of interest to researchers in systems biology and multi-omics integration, but its impact on advancing understanding of disease mechanisms is likely to be moderate.

    3. Author response:

      General Statements:

      We appreciate the reviewers for the critical review of the manuscript and the valuable comments. We have carefully considered the reviewer’s comments and have revised our manuscript accordingly.

      Point-by-point description of the revisions:

      Reviewer #1 (Evidence, reproducibility and clarity):

      Major comments

      (1) This study leaves out lipid metabolism as a major energy metabolism pathway relevant to AD. The authors themselves cite the significance of acylcarnitines and CPT1A in AD (pg. 3, lines 32-33, pg. 4, lines 1-2). Lipid metabolism and homeostasis is known to be disrupted in AD1. Fatty acid oxidation is a known energy source in the prefrontal cortex2 and will also generate acetyl coA, which this study reveals is a significant decreased metabolite in AD. Furthermore, sphingomyelin emerges as one of the major decreased DEMs as well. Thus, lipid metabolism should be highlighted in Figure 3 and discussed throughout the manuscript; otherwise its omission should be clearly stated and justified.

      We appreciate the reviewer’s insightful comment regarding a critical role of lipid metabolism in AD. We recognize that lipid metabolism is a metabolic pathway deeply involved in AD pathology (Baloni et al., 2022, 2020; Varma et al., 2021). Accordingly, we have revised the Limitations section to more strongly emphasize its role as a vital energy source (pg. 13, lines 15-17). Regarding the visualization of lipid metabolism, we extracted lipid-related pathway from the trans-omic network but found that the regulatory relationships among DEPs and DEMs were excessively complex and interconnected. Thus, interpreting this regulatory network seemed to be more challenging compared to the other energy production pathways presented in our manuscript. Therefore, we have concluded that the pathway analysis in our trans-omic network may not be suitable for deeply elucidating the lipid dysregulation in AD. We have added a statement acknowledging this as a limitation of our current methodology in the revised manuscript (pg. 13, lines 13-22).

      (2) The covariates used for differential analysis should be discussed and justified. Notably, age is used as a covariate for transcriptomic analysis but not proteomic and metabolomic analysis, with no justification. Additionally, given the known importance of lipid metabolism in AD and the putative role of APOE in lipid homeostasis3, APOE genetic status should be considered as a covariate, or its omission should be justified.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, such as age at death and RIN, is that these data were not available for each sample. Thus, we referred to the original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, and years of education. Regarding the proteomic dataset, in the original article (Johnson et al., 2020), age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (3) The authors make a conclusion statement that suggests intervention: "Collectively, our data suggests that preserving or improving the ability to produce ATP and early intervention in the process of nitrogen metabolism are candidates for the prevention and treatment of dementia" (pg. 12, lines 12-14). This claim is not well-supported by the evidence provided in the study. There are a few limitations: (a) This was an observational, not interventional study; (b) The study did not establish whether the metabolic disruptions are causes or effects in AD; and (c) ATP or other bioenergetic indicators were not directly measured. Therefore, any statements about potential interventions should be removed or qualified as highly speculative.

      We agree with the reviewer that the statement regarding potential interventions was not sufficiently supported by our analyses. Accordingly, we have removed the sentence regarding prevention and treatment from the revised manuscript (e.g., we have deleted final paragraph of the previous manuscript).

      (4) In conjunction with the last point, the main conclusion of the study is that energy production is down in AD. The data presented in Figure 3 are consistent with this conclusion, but it is far from definitive due to limitations stated above in comments 3a and 3b. The authors should offer additional support for this conclusion: experimental follow-up, flux modeling, analysis of alternative datasets with ATP measurement, causal inference.

      We sincerely thank the reviewer for this valuable and constructive suggestion. Regarding flux modeling, we agree that metabolic flux analysis could provide important mechanistic insight. Indeed, previous studies have applied flux modeling in the context of lipid metabolism in Alzheimer’s disease (Baloni et al., 2022). We also attempted to perform flux modeling focusing on energy metabolism. However, we found it difficult to obtain biologically meaningful and robust results and therefore decided not to include these analyses in the current manuscript.

      With respect to ATP measurements, we fully agree that direct evidence of altered ATP levels would further strengthen our conclusion. However, to the best of our knowledge, there are currently no publicly available large-scale datasets that directly measure ATP levels in human postmortem brain tissues. This limitation makes it challenging to incorporate validation in the present study.

      Regarding experimental follow-up, we agree that functional validation is essential to confirm the mechanistic implications of our findings. We are actively considering follow-up experimental studies. However, we consider the present work to be a multi-omic integrative analysis aimed at identifying key molecular alterations and generating biologically important hypotheses. We have revised the Limitation section to more clearly position this manuscript as an observational systems-level analysis (pg. 13, lines 20-22).

      (5) The validation analysis did not sufficiently show the generalizability of this study's results. The authors demonstrated a correlation of 0.53 to the MSBB transcriptomics data and 0.60 to the AMP-AD DiverseCohorts proteomics data. Beyond these correlation coefficients, no meaningful comparison between the datasets is offered. How concordant are the differentially expressed features (or pathways) between the datasets? How robust would the trans-omic network be if incorporating the alternate datasets? Is the main conclusion (energy metabolism is down in AD) supported by the validation datasets? We think this analysis should be expanded and described in the main text.

      Although the results for external metabolomics datasets are reported in Fig S2C, correlation coefficients with the external data are not reported. The authors state, "Note that each study used different definitions for AD and CT groups, had variations in measurement methods and brain regions analyzed." We appreciate these limitations. However, the external data should be re-analyzed using the same definitions of AD and CT, if possible. The limitations and results (which DEMs are shared between datasets) should be discussed in the main text.

      We thank the reviewer for this important comment regarding the generalizability of our findings. In the revised manuscript, we have expanded the validation analyses and summarized the results in Figure S2. First, at the transcriptomic level, Figure S2B and S2C show the overlap between up- and downregulated genes in AD identified in our ROSMAP-derived analyses and those reported in a previously published large-scale meta-analysis of 2,114 postmortem samples across seven brain regions (Wan et al., 2020). A substantial proportion of DEGs were shared, supporting cross-cohort and cross-region robustness to some extent. At the proteomic level, Figure S2E shows a comparison between the ROSMAP and the AMP-AD DiverseCohorts datasets. We highlighted the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3 and calculated a separate correlation coefficient for this subset (Pearson coefficient = 0.86, p-value = 1.5e-7), further supporting our main conclusion. In addition, to assess the concordance between the two datasets in a threshold-independent manner, we additionally performed Rank-Rank Hypergeometric Overlap (RRHO) analysis (Figure S2E). RRHO analysis (Cahill et al., 2018; Plaisier et al., 2010) enables the comparison of ranked protein lists without relying on arbitrary differential expression cutoffs and has been used for cross-dataset comparison in several previous studies (Fröhlich et al., 2024; Maitra et al., 2023). The RRHO heatmaps demonstrated significant enrichment in the concordant quadrants, confirming systematic agreement between datasets beyond simple correlation coefficients. For metabolomics, Figure S2G shows RRHO analyses comparing the ROSMAP metabolomic data with other datasets measured by the same UPLC-MS/MS platform (Batra et al., 2024; Novotny et al., 2023), demonstrating significant concordance in ranked metabolite changes in AD.

      (6) The glycolysis analysis and discussion needs more development. Glycolysis and gluconeogenesis share many of the same enzymes, but they are not the same pathway and should not be discussed as such. To make a claim about the overall influence of enzyme and metabolite levels on glycolysis, the authors should focus on the energetically committing steps of glycolysis (hexokinase, phosphofructokinase, pyruvate kinase) in Figure 3A, and include the full/current version of the figure in the supplement. Gluconeogenesis-specific enzymes (pyruvate carboxylase, PEPCK) are not mentioned at all - are they among the DEPs/DEGs?

      We appreciate the reviewer’s comment regarding the distinction between glycolysis and gluconeogenesis pathway. Among the gluconeogenesis-specific enzyme proteins, G6PC1, FBP1, PC, and PCK2 were measured in our dataset, but none of them were identified as DEPs. In addition, gluconeogenesis is a process that occurs primarily in the liver and kidney rather than the brain. Given this biological context and the lack of significant changes in relevant enzymes, we have revised the terminology throughout the manuscript, replacing “glycolysis/gluconeogenesis pathway” with “glycolysis pathway” in the revised version.

      (7) Given that there wasn't good concordance between the DEGs and DEPs, did including the mRNA and transcription factor layers in the network really add anything useful? It seems like the main conclusions of the manuscript were driven by the protein and metabolite layers only. How many of the DE metabolic enzymes were coregulated at the transcript and protein level? It would be useful to include the 5-layer trans-omic network in the supplement to display these results. Given your network, at what level does it appear that energy metabolism is regulated?

      It is true that our primary conclusion regarding the regulation of energy metabolism is driven by the changes in protein and metabolite abundance. However, we consider the low concordance between mRNA and protein expression itself to be an important feature of AD pathology, as also reported in previous studies (Johnson et al., 2022; Tasaki et al., 2022). Although we did not perform a further analysis of this discordance, we believe that including the TF and mRNA layers into the metabolic trans-omic network strengthens a system-wide view of metabolic dysregulation in AD.

      Regarding the mRNA changes corresponding to the DEP enzymes, please refer to Figure S7A.

      (8) Comment further on the results from Figure 2D. What can be learned from identifying metabolites with the greatest degree centrality? What pathways other than energy metabolism are highlighted by the trans-omic network?

      We assume that some energetic indicators, including AMP and acetyl-CoA, and nitrogen metabolism-related metabolites, Glu, 2-oxoglutarate, and urea, can be potential key regulators of dysregulated metabolism in AD.

      (9) (Suggestion) We suggest the authors leverage their trans-omic network in additional ways beyond giving a snapshot of a few energy metabolism pathways. The analysis of top DEMs could go further. What pathways are impacted beyond energy metabolism? Among the metabolic reactions allosterically regulated by top DEMs, what metabolic pathways are enriched?

      We identified the enriched metabolic pathways that were allosterically regulated by DEMs in AD using Fisher’s exact test. Alanine, aspartate, and glutamate metabolism pathways were significantly enriched in 2-oxoglutarate, glutarate, alanine, and glutamate-regulating metabolic reactions. Arginine and proline metabolism pathway was enriched in N-methyl-L-arginine and putrescine-regulating metabolic reactions. Arginine biosynthesis pathway was enriched in arginine-regulating metabolic reactions. Glycerophospholipid metabolism pathway was enriched in CDP-ethanolamine-regulating metabolic reactions. Glycine, serine, and threonine metabolism pathway was enriched in serine-regulating metabolic reactions. Purine metabolism pathway was enriched in AMP-regulating metabolic reactions. Pyrimidine metabolism pathway was enriched in deoxyuridine and thymidine-regulating metabolic reactions. Sphingolipid metabolism pathway was enriched in sphingosine-regulating metabolic reactions. However, this analysis did not yield sufficiently valuable insights into the regulatory relationships among biomolecules in AD. Thus, we did not include these results in the revised manuscript.

      (10) (Suggestion) Figure 3 shows that most differential signal in AD points to lower energy production due to the combination of differentially expressed metabolites and enzymes, but we are not given much context about the strength of these among all the differential signals. We would suggest including volcano plots where the features of interest, i.e. DE enzymes and metabolites, are colored differently (or a similar figure).

      We thank the reviewer for this constructive suggestion. To provide better context regarding the importance of the differential signals, we have added volcano plots for mRNAs, proteins, and metabolites in Figure S4A, B, and C.

      (11) (Suggestion) The PPI network could be better leveraged to understand metabolic changes in AD. If nodes are grouped into subnetworks (e.g. by Louvain / Leiden clustering) and tested for pathway enrichment, could you find functional subnetworks of coordinately up- and down- regulated metabolic enzymes? This could yield some pathways of interest beyond the energy metabolism pathways already highlighted.

      We appreciate the reviewer’s suggestion to utilize the PPI network for subnetwork analysis. However, it is important to note that the proteomic dataset analyzed in this study is derived from the original work of (Johnson et al., 2020). In that paper, the authors already performed a Weighted Gene Co-expression Network Analysis (WGCNA) across several datasets to identify co-expressed modules and functional pathways.

      Given this, we assumed that applying additional clustering methods to the same dataset would be unlikely to yield significant biological insights beyond the established findings.

      Minor comments

      (1). "All genes" and "all metabolites" should not be the background for the proteomic and metabolic pathway enrichment analysis by Metascape and MetaboAnalyst. The background should be limited to the proteins and metabolites that were measured.

      We fully agree with the reviewer that using “all gene” or “all metabolites” as a background is not suitable for enrichment analyses. As suggested, we have revised the enrichment analyses using the measured proteins and metabolites as a background in both Metascape and MetaboAnalyst (Fig. S4D).

      (2) Highlight the metabolic enzymes in Fig S2B. Calculate a separate correlation coefficient for the enzymes extracted in the energy metabolism analysis from Fig 3.

      We appreciate the reviewer’s suggestion to refine the correlation analysis. As requested, we have revised Fig. S2D to explicitly highlight the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3. We calculated a separate correlation coefficient for the subset (Pearson coefficient = 0.86, p-value = 1.5e-7).

      (3) Use a multiple hypothesis adjusted p-value or q-value in Figure S3.

      We agree with the reviewer regarding the necessity of correcting for multiple comparisons. Accordingly, we have revised Fig. S4D using q-values.

      (4) Describe the methods used to calculate the logFC values from the validation dataset.

      We have revised the Methods to include a detailed description of the procedure used to calculate the log2FC values for the validation datasets (pg. 21, lines 13-15).

      (5) It is difficult to read Figure 3. We would recommend really emphasizing to the reader to refer to Fig S7B as a "key" to this figure. The description of the red/blue arrows and nodes in the methods section (pg. 24, lines 21-36, pg 25, lines 1-4) were also helpful, but very lengthy. We recommend putting an abridged version of this description into the Fig S7 figure legend.

      We appreciate the feedback regarding the readability of Fig. 3. As recommended, we have revised the manuscript to explicitly direct readers to Fig. S8B as an essential “key” for interpreting the network visualization (pg. 8, lines 28). Furthermore, we have added an abridged description of the network elements to the legend of Fig. S8B.

      (6) The S7 figure legend should refer to panels A and B, not E and F.

      We apologize for this oversight. We have corrected the legend of Fig. S8.

      (7) (Suggestion) Are any of the differentially expressed metabolites allosteric regulators of the DE transcription factors? This could be interesting to discuss.

      We appreciate the reviewer’s insightful suggestion about the potential allosteric regulation of the DETFs by DEMs. We conducted an extensive literature search to identify any reports related to this perspective. However, to the best of our knowledge, no such direct interactions have been reported to date.

      Reviewer #1 (Significance):

      The study's strength lies in leveraging three omics modalities across large patient cohorts (n ~ 150-240) to identify coherent signals between transcriptomics, proteomics, and metabolomics in postmortem DLPFC tissue. It was encouraging to see that the main result, showing downregulation for TCA, oxidative phosphorylation, and ketone body metabolism, emerged from consistent signals across both proteomics and metabolomics. This result was consistent with previous findings in other models cited by the author4,5 and other studies 6,7 demonstrating deficiency in energy-producing pathways in AD.

      Another strength of the study is the application of thoughtful methodology to connect differentially expressed proteins and metabolites via an intermediate data layer of metabolic reactions. The authors leverage the KEGG and BRENDA databases and apply sound logic to estimate the effects of enzyme level and metabolite level on pathway activity, with metabolites serving as substrate, product, or allosteric regulator for reactions. This trans-omic network methodology was developed in previous studies cited by the author8,9.

      However, as written, this study is limited in its contribution of new knowledge to the AD research field. The main conclusion (energy production is down in AD, due to regulatory disruption of energy metabolism) is not strongly supported (see comments 1, 3, and 4 for elaboration). The evidence could be improved by orthogonal approaches: further experimentation, further integration of external datasets, causal modeling, or flux modeling. Alternatively, even in the absence of new experimental and computational approaches, the story could be made more complete by further leveraging the trans-omic network to provide insights into (a) the regulation of energy metabolism; and (b) the impacts of key disrupted metabolites (see comments 7-9).

      The study is also limited in its demonstrating the power of these methodologies to provide integrative insights. As mentioned above, the integration of enzyme levels and metabolite levels is clearly useful (Figure 3). In contrast, the utility of the mRNA and transcription factor layers was not evident. The study did not appear to improve or expand upon trans-omic network methodology described in the previous works. Finally, the various analyses (analyzing the trans-omic network for nodes with the highest degree centrality, the PPI analysis, and viewing the energy metabolism pathways in the network) provided disparate results that were only tenuously connected in the discussion section.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary

      This manuscript integrates public transcriptomic, proteomic, and metabolomic datasets from ROSMAP DLPFC samples to construct a multi-layer metabolic trans-omic network in Alzheimer's disease. By linking transcription factors, enzyme mRNAs, proteins, metabolic reactions, and metabolites, the authors report coordinated downregulation of the TCA cycle, oxidative phosphorylation, and ketone body metabolism, along with mixed regulatory signals in glycolysis/gluconeogenesis. They interpret these patterns as indicative of broad energetic dysfunction and alterations in amino-acid/nitrogen metabolism in AD. While the framework is conceptually appealing, much of the analysis remains descriptive, and several biological interpretations extend beyond what the data can robustly support. The reliance on bulk tissue without accounting for cell-type composition, limited covariate adjustment, and the absence of validation or sensitivity analyses reduce confidence in the mechanistic conclusions. Overall, the study provides a preliminary systems-level overview, but additional rigor is needed before the proposed trans-omic regulatory insights can be considered convincing.

      Major Comments

      (1) Interpretation requires more cautious phrasing, and validation is essential. The manuscript frequently asserts that specific pathways are "inhibited" or that energetic deficits are "compensated," but these conclusions extend beyond what the descriptive, bulk-level data can support. Because no metabolic flux, causality, or direct functional measurements are included, the results should be framed as putative regulatory shifts, not confirmed impairments. Critically, key claims about pathway inhibition would require flux modeling, perturbation analyses, or experimental validation to be convincing. Without such validation, the mechanistic interpretations remain speculative.

      We thank the reviewer for this crucial comment. We fully agree that, given the descriptive and bulk-level nature of our analysis, mechanistic interpretations must be made with caution. In the absence of direct metabolic flux measurements or experimental validation, our findings should be interpreted as putative regulatory shifts rather than confirmed functional impairments. Accordingly, we have revised the manuscript to temper mechanistic claims. We have replaced definitive statements with more speculative phrasing (e.g., “Our analysis revealed a putative coordinated downregulation …” instead of “Our analysis revealed a coordinated downregulation …” in Abstract section; “we demonstrate the systems-level view of the potential dysregulated energy production …” instead of “we demonstrate the systems-level view of the dysregulated energy production …” in pg. 10, lines 25-26).

      (2) Although the authors acknowledge this in the limitations, bulk-level differences may primarily reflect altered proportions of neurons, astrocytes, microglia, and oligodendrocytes rather than true within-cell-type regulation. Incorporating a cell-type deconvolution or performing a sensitivity analysis would substantially improve interpretability. This issue also impacts the trans-omic network: if the molecules included originate from different cell types, the inferred regulatory relationships may not reflect true intracellular processes.

      We appreciate the reviewer’s point that bulk-level differences can reflect altered proportions of different brain cell types, subsequently affecting the inferred trans-omic network analysis. To assess the changes in cell type proportions of the samples that we used in our study, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglias, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two group. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      (3) Differential analysis covariates. For the differential expression analyses, only gender and PMI were included as covariates. Additional variables, such as age at death, RIN, neuropathological measures, and comorbidities, can strongly influence molecular profiles and should be considered to ensure that the observed differences reflect AD-related biology rather than confounding pathological or technical factors.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, including age at death and RIN, is that these data for each sample were not available. Thus, we referred to original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, or education. Regarding the proteomic dataset, in the original article, age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (4) Network stability and sample non-overlap. Proteomic, transcriptomic, and metabolomic data come from partially overlapping individuals. The authors should test whether the reconstructed network is robust to: different significance thresholds, restricting analyses to overlapping samples and alternative definitions of AD vs control.

      We appreciate the reviewer’s comment for the trans-omic network stability. In our study, the number of individuals for whom all omic modalities were measured was relatively small (n=25 in CT and n=35 in AD). This limited overlap reduces statistical power and can affect the downstream network construction. We have acknowledged this limitation in the revised manuscript and clarified that the reconstructed networks should be interpreted with caution regarding reproducibility and generalizability (pg. 13, lines 13-23).

      Minor Comments

      (1) Some TF enrichment and regulatory inferences lack explicit mention of multiple-testing correction.

      We apologize for the lack of clarity in our original description. We have corrected for multiple-testing for the TF inference. Thus, we have revised the Methods section to explicitly describe the correction method used and the threshold applied (pg. 23, lines 23-24).

      (2) The limitations section is strong but should explicitly discuss the influence of postmortem interval on metabolite levels.

      We appreciate the reviewer’s comment about the effect of postmortem interval on changes in metabolite levels. Accordingly, we have added the description of this perspective in our revised manuscript (pg. 13, lines 1-5).

      Reviewer #2 (Significance):

      The study extends a trans-omic integration framework, originally applied to metabolic disease, into the context of Alzheimer's pathology. Although the biological findings largely confirm known alterations in mitochondrial and energy metabolism, the network-based approach offers a structured way to view cross-layer regulatory changes. Its main advance is conceptual rather than biological, providing a unified framework rather than uncovering fundamentally new mechanisms. This work will primarily interest researchers in neurodegeneration and systems biology, as well as computational groups developing multi-omics integration methods.

      Reviewer #3 (Evidence, reproducibility and clarity):

      This study leverages existing transcriptomic, metabalomic and proteomic datasets from prefrontal cortex (PFC) to assess metabolic dysregulation in Alzheimer's disease (AD). They found a downregulation of multiple metabolic pathways, including TCA cycle, oxidative phosphorylation, and ketone metabolism, that may explain bioenergetic alterations in AD.

      The study used matching ROSMAP omics datasets from the DLPFC that have allowed more robust data integration. However, the datasets are all generated using bulk tissue, which makes data interpretation difficult. For example, the AD changes they observed may be due to shifts in cell type proportion with disease (e.g. cell death, neuron inflammation). Did the authors account for any potential shifts in cell type proportion in their analysis?

      If the assumption is that the changes in AD are cell intrinsic, which cell types are likely to be impacted? Can the authors integrate any existing single-cell analysis to infer which cell types may be driving the signals they detect, and whether this accounts for some of the antagonistic regulatory effects that were detected?

      We thank the reviewer for their insightful comments. We agree that the use of bulk tissue datasets cannot account for cell-type heterogeneity. As noted in our Limitations section (pg. 12, lines 24-27), we recognize that previous studies have found that the Braak stage is correlated positively with microglia and astrocyte proportions and negatively with oligodendrocyte proportion (Hannon et al., 2024; Shireby et al., 2022). Regarding the integration of single-cell analysis, we have referenced recent snRNA-seq findings (Mathys et al., 2024) in our Limitations section (pg. 12, lines 28-32) to deconvolve our bulk signatures.

      Furthermore, in our revised manuscript, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglia, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two groups. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      Reviewer #3 (Significance):

      The manuscript provides multimodal insight into metabolic dysregulation in AD in the PFC. Given that metabolic dysfunction is likely to play a major in disease pathogenesis, this is a study of importance. However, the findings lack granularity at the cell type level, which limits the impact of the study.

      Reference

      (1) Baloni, P., Arnold, M., Buitrago, L., Nho, K., Moreno, H., Huynh, K., Brauner, B., Louie, G., Kueider-Paisley, A., Suhre, K., Saykin, A. J., Ekroos, K., Meikle, P. J., Hood, L., Price, N. D., Alzheimer’s Disease Metabolomics Consortium, Doraiswamy, P. M., Funk, C. C., Hernández, A. I., … Kaddurah-Daouk, R. (2022). Multi-Omic analyses characterize the ceramide/sphingomyelin pathway as a therapeutic target in Alzheimer’s disease. Communications Biology, 5(1), 1074.

      (2) Baloni, P., Funk, C. C., Yan, J., Yurkovich, J. T., Kueider-Paisley, A., Nho, K., Heinken, A., Jia, W., Mahmoudiandehkordi, S., Louie, G., Saykin, A. J., Arnold, M., Kastenmüller, G., Griffiths, W. J., Thiele, I., Alzheimer’s Disease Metabolomics Consortium, Kaddurah-Daouk, R., & Price, N. D. (2020). Metabolic Network Analysis Reveals Altered Bile Acid Synthesis and Metabolism in Alzheimer’s Disease. Cell Reports. Medicine, 1(8), 100138.

      (3) Batra, R., Arnold, M., Wörheide, M. A., Allen, M., Wang, X., Blach, C., Levey, A. I., Seyfried, N. T., Ertekin-Taner, N., Bennett, D. A., Kastenmüller, G., Kaddurah-Daouk, R. F., Krumsiek, J., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2023). The landscape of metabolic brain alterations in Alzheimer’s disease. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(3), 980–998.

      (4) Batra, R., Krumsiek, J., Wang, X., Allen, M., Blach, C., Kastenmüller, G., Arnold, M., Ertekin-Taner, N., Kaddurah-Daouk, R., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2024). Comparative brain metabolomics reveals shared and distinct metabolic alterations in Alzheimer’s disease and progressive supranuclear palsy. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 20(12), 8294–8307.

      (5) Cahill, K. M., Huo, Z., Tseng, G. C., Logan, R. W., & Seney, M. L. (2018). Improved identification of concordant and discordant gene expression signatures using an updated rank-rank hypergeometric overlap approach. Scientific Reports, 8(1), 9588.

      (6) Fröhlich, A. S., Gerstner, N., Gagliardi, M., Ködel, M., Yusupov, N., Matosin, N., Czamara, D., Sauer, S., Roeh, S., Murek, V., Chatzinakos, C., Daskalakis, N. P., Knauer-Arloth, J., Ziller, M. J., & Binder, E. B. (2024). Single-nucleus transcriptomic profiling of human orbitofrontal cortex reveals convergent effects of aging and psychiatric disease. Nature Neuroscience, 27(10), 2021–2032.

      (7) Green, G. S., Fujita, M., Yang, H.-S., Taga, M., Cain, A., McCabe, C., Comandante-Lou, N., White, C. C., Schmidtner, A. K., Zeng, L., Sigalov, A., Wang, Y., Regev, A., Klein, H.-U., Menon, V., Bennett, D. A., Habib, N., & De Jager, P. L. (2024). Cellular communities reveal trajectories of brain ageing and Alzheimer’s disease. Nature, 633(8030), 634–645.

      (8) Hannon, E., Dempster, E. L., Davies, J. P., Chioza, B., Blake, G. E. T., Burrage, J., Policicchio, S., Franklin, A., Walker, E. M., Bamford, R. A., Schalkwyk, L. C., & Mill, J. (2024). Quantifying the proportion of different cell types in the human cortex using DNA methylation profiles. BMC Biology, 22(1), 17.

      (9) Johnson, E. C. B., Carter, E. K., Dammer, E. B., Duong, D. M., Gerasimov, E. S., Liu, Y., Liu, J., Betarbet, R., Ping, L., Yin, L., Serrano, G. E., Beach, T. G., Peng, J., De Jager, P. L., Haroutunian, V., Zhang, B., Gaiteri, C., Bennett, D. A., Gearing, M., … Seyfried, N. T. (2022). Large-scale deep multi-layer analysis of Alzheimer’s disease brain reveals strong proteomic disease-related changes not observed at the RNA level. Nature Neuroscience, 25(2), 213–225.

      (10) Johnson, E. C. B., Dammer, E. B., Duong, D. M., Ping, L., Zhou, M., Yin, L., Higginbotham, L. A., Guajardo, A., White, B., Troncoso, J. C., Thambisetty, M., Montine, T. J., Lee, E. B., Trojanowski, J. Q., Beach, T. G., Reiman, E. M., Haroutunian, V., Wang, M., Schadt, E., … Seyfried, N. T. (2020). Large-scale proteomic analysis of Alzheimer’s disease brain and cerebrospinal fluid reveals early changes in energy metabolism associated with microglia and astrocyte activation. Nature Medicine, 26(5), 769–780.

      (11) Maitra, M., Mitsuhashi, H., Rahimian, R., Chawla, A., Yang, J., Fiori, L. M., Davoli, M. A., Perlman, K., Aouabed, Z., Mash, D. C., Suderman, M., Mechawar, N., Turecki, G., & Nagy, C. (2023). Cell type specific transcriptomic differences in depression show similar patterns between males and females but implicate distinct cell types and genes. Nature Communications, 14(1), 2912.

      (12) Mathys, H., Boix, C. A., Akay, L. A., Xia, Z., Davila-Velderrain, J., Ng, A. P., Jiang, X., Abdelhady, G., Galani, K., Mantero, J., Band, N., James, B. T., Babu, S., Galiana-Melendez, F., Louderback, K., Prokopenko, D., Tanzi, R. E., Bennett, D. A., Tsai, L.-H., & Kellis, M. (2024). Single-cell multiregion dissection of Alzheimer’s disease. Nature, 632(8026), 858–868.

      (13) Novotny, B. C., Fernandez, M. V., Wang, C., Budde, J. P., Bergmann, K., Eteleeb, A. M., Bradley, J., Webster, C., Ebl, C., Norton, J., Gentsch, J., Dube, U., Wang, F., Morris, J. C., Bateman, R. J., Perrin, R. J., McDade, E., Xiong, C., Chhatwal, J., … Harari, O. (2023). Metabolomic and lipidomic signatures in autosomal dominant and late-onset Alzheimer’s disease brains. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(5), 1785–1799.

      (14) Plaisier, S. B., Taschereau, R., Wong, J. A., & Graeber, T. G. (2010). Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures. Nucleic Acids Research, 38(17), e169.

      (15) Shireby, G., Dempster, E. L., Policicchio, S., Smith, R. G., Pishva, E., Chioza, B., Davies, J. P., Burrage, J., Lunnon, K., Seiler Vellame, D., Love, S., Thomas, A., Brookes, K., Morgan, K., Francis, P., Hannon, E., & Mill, J. (2022). DNA methylation signatures of Alzheimer’s disease neuropathology in the cortex are primarily driven by variation in non-neuronal cell-types. Nature Communications, 13(1), 5620.

      (16) Tasaki, S., Xu, J., Avey, D. R., Johnson, L., Petyuk, V. A., Dawe, R. J., Bennett, D. A., Wang, Y., & Gaiteri, C. (2022). Inferring protein expression changes from mRNA in Alzheimer’s dementia using deep neural networks. Nature Communications, 13(1), 655.

      (17) Varma, V. R., Wang, Y., An, Y., Varma, S., Bilgel, M., Doshi, J., Legido-Quigley, C., Delgado, J. C., Oommen, A. M., Roberts, J. A., Wong, D. F., Davatzikos, C., Resnick, S. M., Troncoso, J. C., Pletnikova, O., O’Brien, R., Hak, E., Baak, B. N., Pfeiffer, R., … Thambisetty, M. (2021). Bile acid synthesis, modulation, and dementia: A metabolomic, transcriptomic, and pharmacoepidemiologic study. PLoS Medicine, 18(5), e1003615.

      (18) Wan, Y.-W., Al-Ouran, R., Mangleburg, C. G., Perumal, T. M., Lee, T. V., Allison, K., Swarup, V., Funk, C. C., Gaiteri, C., Allen, M., Wang, M., Neuner, S. M., Kaczorowski, C. C., Philip, V. M., Howell, G. R., Martini-Stoica, H., Zheng, H., Mei, H., Zhong, X., … Logsdon, B. A. (2020). Meta-Analysis of the Alzheimer’s Disease Human Brain Transcriptome and Functional Dissection in Mouse Models. Cell Reports, 32(2), 107908.

    1. eLife Assessment

      The authors propose a pipeline using large language models (LLMs) to benchmark experimentally well validated computational models of signaling networks. The findings are important, with available methods that guide testing future models and find new molecular interactions. Furthermore, the work shows how general-purpose LLMs can generate up to 91% of reactions of a bacterial metabolic network. The support is convincing and offers a number of performance metrics.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors did an excellent job with their resubmission, politely and elegantly answering the comments from the reviewers.]

      Summary:

      Large language models (LLMs) have been developed rapidly in recent years and are already contributing to progress across scientific fields. The manuscript tries to address a specific question: whether LLMs can accurately infer signaling networks from gene lists.

      Strengths:

      The manuscript raises a good question: whether current LLMs can accurately generate signaling networks from gene lists.

    3. Reviewer #2 (Public review):

      Summary:

      The authors evaluate whether commonly used LLMs (ChatGPT, Claude and Gemini) can reconstruct signalling networks and predict effects of network perturbations, and propose a pipeline for benchmarking future models. Across three phenotypes (hypertrophy, fibroblast signalling, and mechanosignalling), LLMs capture upstream ligand-receptor interactions and conserved crosstalk but fail to recover downstream transcriptional programmes. Logic-based simulations show that LLM-derived networks underperform compared to manually curated models. The authors also propose that their pipeline can be used for benchmarking future models aimed at reconstructing signalling networks.

      Strengths:

      The authors compare the outcomes from three LLMs with three manually curated and validated models. Additionally, they have investigated gene network reconstruction in the context of three distinct phenotypes. Using logic-based modelling, the authors assessed how LLM-derived networks predict perturbation effects, providing functional validation beyond network overlap.

      Weaknesses:

      The authors have used legacy models for all three LLMs, and the study would benefit from testing the current versions of the LLMs (ChatGPT 5.2, Claude 4.5 and Gemini 2.5). Additional metrics such as node coverage, node invention, direction accuracy and sign accuracy would be useful to make robust comparisons across models.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary

      Large language models (LLMs) have been developed rapidly in recent years and are already contributing to progress across scientific fields. The manuscript tries to address a specific question: whether LLMs can accurately infer signaling networks from gene lists. However, the evaluation is inadequate due to four major weaknesses described below. Despite these limitations, the authors conclude that current general-purpose LLMs lack adequate accuracy, which is already widely recognized. Its key contribution should instead be to provide concrete recommendations for the development of specialized LLMs for this task, which is completely absent. Developing such specific LLMs would be highly valuable, as they could substantially reduce the time required by researchers to analyze signaling networks.

      Strengths

      The manuscript raises a good question: whether current LLMs can accurately generate signaling networks from gene lists.

      Weaknesses:

      (1) The authors evaluate LLM performance using only three signaling networks: "hypertrophy", "fibroblast", and "mechanosignaling". Given the large number of well-established signaling pathways available, this is not a comprehensive assessment. Moreover, the analysis need not be restricted to signaling networks. Other network types, including metabolic and transcriptional regulatory networks, are already accessible in well-known databases such as KEGG, Reactome, BioCyc, WikiPathways, and Pathway Commons. Including these additional networks would substantially strengthen the evaluation.

      We agree with the reviewer that our evaluation of LLM performance is not comprehensive of all signaling networks, and that the benchmarking was previously limited to signaling networks. The purpose of this study is to benchmark LLMs against peer-reviewed computational models that make testable predictions and are highly validated experimentally, of which these three signaling networks are strong examples. KEGG, Reactome, Biocyc, WikiPathways, and Pathway Commons are databases that contain collections of individual reactions or pathways but are not computational models themselves and have not been experimentally validated in that sense.

      While this study focuses primarily on signaling networks, we agree that it would be useful to evaluate how well LLMs perform in generating networks of another type, for which predictive and validated computational models are available. Therefore, in new Figure 3 we now test the ability of LLMs to reconstruct the E. Coli core metabolic network, as well as test its ability to predict growth on metabolic substrates using flux balance analysis. We find that Claude Opus 4.6, GPT 5.2 Pro, and Gemini 3 Pro Preview perform well at reconstructing reactions from E. Coli core metabolism, but these reactions are not sufficient to predict growth on a variety of substrates. Given this expansion of scope, we replaced “signaling networks” in the title with “biochemical networks”.

      (2) In LLM evaluation, the authors use the gene lists that exactly match those in their "ground truth" networks, thereby fixing the set of nodes and evaluating only the predicted edges. However, in practical research, the relevant genes or nodes are not fully known. A more realistic assessment would therefore include gene lists with both genes present in the ground-truth network and additional genes absent from it, to evaluate the ability of the LLM to exclude irrelevant genes.

      We agree with the reviewer that evaluating the capacity of these LLMs to exclude additional genes is interesting. But because biological networks are always incompletely known, there is no “ground truth” of genes absent from a given network. Therefore, for the most rigorous benchmarking against a “ground truth”, we examine the positive predictive ability of LLMs. However, in response to this comment and point 3 below, we further examine additional measures of performance that include “false positives”.

      (3) The authors report only the recall/sensitivity of the LLM, without assessing specificity. In practical applications, if an LLM generates a large number of incorrect interactions that greatly exceed the correct ones, researchers may be misled or may lose confidence in the LLM output. Therefore, a comprehensive evaluation must include both sensitivity and specificity. Furthermore, it would be informative to check whether some of the "false positives" might in fact represent biologically plausible interactions that are absent from the manually curated "ground truth". Manually generated "ground truth" can overlook genuine interactions, and the ability of LLMs to recover such missing edges could be particularly valuable. This may even represent one of the most important potential contributions of LLMs.

      We agree with the reviewer that additional metrics could inform the evaluation of LLM performance. Therefore, as recommended, we calculated sensitivity, specificity, precision, negative predictive value (NPV), accuracy, and F1 score for each of the network models (hypertrophy, fibroblast, and mechanosignaling). These new results are summarized in confusion matrices shown in a new Supplementary Figures 3, 4, and 5. We performed this additional benchmarking across the 10 replicates for each LLM.

      One limitation of this approach is the substantial class imbalance within these confusion matrices. Because we interpret actual negatives as connections that are not found between any nodes of the ground truth models, there will be >10x more true negatives than any of the other classes. This makes specificity, accuracy, and the negative predictive value less informative.

      To illustrate this point, consider the ground-truth hypertrophy network which contains 191 connections between 106 nodes. The total number of possible connections between any two nodes is 106<sup>2</sup> = 11,236. Given that there are 191 actual positives, that leaves 11,045, actual negatives (as illustrated in the null predictor of Supplementary Figure 3A). As the number of node-to-node connections predicted by LLMs is on the order of a few hundred, the number of true negatives is always in the thousands, often outweighing the true positives in the specificity or accuracy calculations.

      The effects of this class imbalance are illustrated with the results of the “null predictors” in Supplementary Figure 4A, which are hypothetical models that fail to predict any connection between nodes (have predicted positive values of 0). These null predictors have sensitivities of 0 and specificities of 1 and high accuracies and NPVs because of the high true negative rates.

      The precision and F1 scores calculated using these confusion matrices are robust to these class imbalances. Indeed, there is substantial heterogeneity in the number of “false positives” connections generated by the LLMs as illustrated by the precision and F1 scores. We are hesitant to unequivocally label these novel, predicted connections as false positives because, as the reviewer points out, these connections could represent true molecular interactions that were undiscovered at the time of the ground truth models’ conception but have since been demonstrated experimentally and published.

      (4) It is widely known that applying differential equation models to highly complex biological networks, such as the three networks in the manuscript, is meaningless, because these systems involve a large number of parameters whose values can drastically alter the results. As Richard Feynman once said: "with four parameters I can fit an elephant, and with five I can make him wiggle his trunk." Thus, the evaluation of LLMs on "logic-based differential equation models" does not make much sense.

      Differential equation models have been the primary mathematical framework for studying complex systems for decades. We refer readers unfamiliar with differential equations to the Nobel prize-winning work of Hodgkin/Huxley (Physiology or Medicine1963, action potential of neurons), Prigogine (Chemistry 1977, non-equilibrium thermodynamics and pattern formation), John Nash (Economics 1994, game theory and Nash Equilibrium), Merton/Scholes (Economics 1997, dynamics of financial derivatives), and Manabe/Hasselmann (Physics 2021, dynamic modeling of atmosphere and oceans).

      We are confused by the quote of a joke by Richard Feynman about fitting equations to data in the shape of an elephant. While this famous joke is amusing, it is both misattributed (it was a recollection by Enrico Fermi in 1953 of a joke once made by John von Neumann) and deliberately hyperbolic (see https://en.wikipedia.org/wiki/Von_Neumann%27s_elephant). Regardless, the relevance of this joke to our study is unclear, because we are not fitting equations to data. As described in the text, in previous studies we validated the predictions of these three logic-based network models with experimental data that was not used to develop the models.

      Reviewer #1 (Recommendations for the authors):

      (1) All figures are in very poor resolution.

      Thank you for identifying this. We have fixed this issue, which was due to SVG embedding. We now embed as higher resolution PNG and provide full resolution files separately.

      (2) The manuscript does not include data availability or code availability.

      As described in the Methods, all code and data is now available via GitHub (https://github.com/saucermanlab/LLM-network-generation).

      Reviewer #2 (Public review):

      (1) Information on the accuracy of directionality of interaction would help understand if there is a bias towards either a positive or negative association.

      To assess if LLMs are biased in predicting either stimulatory or inhibitory connections, we examined the proportion of stimulatory and inhibitory connections for each ground truth model along with the prediction sets from the different LLMs (Author response table 1). These findings suggest that any directional bias is minimal and not conserved across the different ground truth models. These tables were not included in the revised manuscript.

      Author response table 1.

      Proportion of stimulatory and inhibitory connections present in each ground truth model and in the sets of connections predicted by each LLM (Claude, GPT, and Gemini).

      The primary subset of reaction types that the LLM’s tend to miss are often downstream, cell-type specific, and/or gene regulatory connection (e.g. Figure 1B). This observation is consistent with the fact that the ground truth models were constructed using experimental evidence from specific publications involving defined experimental models and cell types whereas the LLM’s presumably draw from the entire corpus of published literature.

      (2) Do all LLMs capture similar information, or are some LLMs better at capturing certain information than others? Further to this, it would be interesting to look into whether amalgamating information across all three LLMs results in a more accurate network.

      The reviewer asks interesting questions that can be qualitatively answered in Figure 1B and in the network visualizations (Supplementary Figures 1 and 2). These diagrams illustrate predicted connections that are shared between the nodes. To include a more quantitative, comprehensive assessment of this overlap, we have included Venn diagrams showing the extent to which LLMs capture shared information (Supplementary Figures 3-5). One such Venn diagram for the hypertrophy model is included in Supplementary Figure 3C. Indeed, it seems that there is substantial overlap in the “false positive” connections predicted by Claude and GPT. It could be interesting to evaluate if these connections are reflective of newly discovered molecular interactions, as referenced in our response to reviewer one point 3.

      While amalgamating information across all three LLMs might generate a more accurate network, these Venn diagrams illustrate that there remain connections within the ground truth models that are not represented in any of the prediction sets from the LLMs. Indeed, the union of all predicted connections for the hypertrophy network made by any of the 10 replicates from the different LLMs would still lack 33 ground truth connections (Supplemental Figure 3C).

      (3) Would it be possible to retrieve a confidence value of the interactions from the LLMs and conduct Precision, recall rate, AUPR and calibration analyses? These metrics would also help with the benchmarking process.

      We agree with the reviewer that including additional evaluative metrics is instructive. Therefore, for each reference network, we have included precision, specificity, recall, accuracy, and F1 score calculations (Supplementary Figures 3-5). Additionally, we conducted calibration analyses to illustrate the LLM’s reliability. While this latter analysis is interesting, we note that the calibration curves were generated using the frequency of predictions among the 10 replicates as a proxy for confidence. This frequency is highly sensitive to the temperature parameter of the LLMs. Different temperature settings could substantially influence the trajectories of these confidence curves and the distributions of the associated histograms. See Supplemental Figure 3B for calibration analyses of the hypertrophy model.

      We considered performing analyses resembling AUPR as suggested by the reviewer, but we did see a defensible way to vary a “threshold” for the classification of a connection as positive or negative. In this study, a connection predicted by the LLM either does or does not exist within the ground truth mode.

      Typos:

      (1) Generated networks to predict THE CLASSIC "fetal gene program" gene expression.

      Thank you for catching this error. We have corrected it.

      (2) A manually curated network has A functional accuracy of

      Thank you for catching this error. We have corrected it.

    1. eLife Assessment

      The study presents important findings revealing previously unresolved conformational dynamics of the heterodimeric type IV ABC transporter TmrAB using single-molecule FRET. The evidence presented is convincing, integrating careful experimental design with computational approaches to uncover states that are typically masked and difficult to detect. The work will be of interest to scientists studying the molecular mechanisms of primary active transport processes.

    2. Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

    3. Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATP-bound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATPbound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I mentioned that the final section of the Results part seems like an afterthought, especially since the heading suggests a broader scope.

      Reply: We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      The changes made to the section do not align with the wording of the reply. Please consider modifying it further.

      We appreciate the positive feedback and this final comment. We have revised the final section of the Results to better reflect the scope indicated by the heading. In addition to clarifying the kinetic analysis, we now explicitly relate our kinetic observations to previously determined thermodynamic measurements, showing that the rapid interconversion of ATP-bound conformations observed during steady-state turnover is consistent with a thermodynamic landscape characterized by a near-zero free-energy difference and entropy–enthalpy compensation. This revision more clearly integrates the kinetic and thermodynamic aspects of the transporter cycle.

    1. eLife Assessment

      This important study addresses a classic debate in visual processing, using a strong method applied to an impressive dataset obtained from a rare clinical population to evaluate hierarchical models of visual object perception. The paper provides compelling evidence that the hierarchical model is only partly supported: as expected, neural responses in ventral visual cortex show increased representational selectivity for faces along the posterior-anterior axes, but the onsets of the signals do not show a temporal hierarchy, indicating more parallel processing.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      Comments on revised version:

      The authors have performed additional analyses and checks and I would now rate the findings as important and compelling.

    3. Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The work is worth publishing for the data alone, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways, and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided, using resampling that preserve dependencies in a hierarchical manner, which is rare. Two types of analyses are also provided to support evidence in favour of the lack of practical significance for some of the comparisons. Collectively, the various analyses and rich illustrations provide convincing evidence in favour of the conclusions.

      Comments on revised version:

      The authors have addressed all my previous comments and the work is mostly limited by the lack of pre-registration and the exploratory nature of some of the analyses. However, with data and code available, other teams can assess the impact of researchers' degrees of freedom on the main outcomes.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      We would like to thank this reviewer for their positive comments and evaluation of our manuscript.

      Weaknesses:

      (1) A control analysis where areas have known differences in response onset should be performed to increase confidence that the proposed analyses would reveal expected results when a difference in response onset was present across areas. From Figure 3, it can be seen that many electrodes are placed in earlier visual areas (V1-V3) that have previously been shown to have earlier broadband responses to visual images compared to VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018). The same analyses as in Figures 4 and 5 should be used comparing VOTC to early visual areas to confirm that the analyses would detect that V1-V3 have earlier onsets compared to VOTC.

      First, we would like to mention that the analyses performed in our paper are commonly accepted analyses to extract time-domain information.

      Yet, the reviewer is right that, considering our claim, evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies would provide further support for our claims. The solution proposed by the reviewer is interesting but the number of face-selective recording contacts/sites in posterior occipital cortex (colored disks in figure 3) is too small for any meaningful comparison. Moreover, while absolute responses to visual images should indeed emerge earlier in early visual cortex than association cortex, it may not be the case for category-selective responses to visual images (which is what our claim is about).

      To address the reviewer’s concern, we used face-selective responses from regions known to have different response onset latencies: occipital and posterior temporal lobe vs. medial temporal lobe structures (e.g., Mormann et al., 2008, https://pmc.ncbi.nlm.nih.gov/articles/PMC2676868/). Waveforms and onset latencies for these regions (OCC, PTL, MTL) are shown in Author response image 1. Despite the small number of contacts showing significant face-selective activity in the MTL (N=20) and the lower SNR in this region, the onset latency differences between OCC/PTL and MTL are significant using all 4 methods of latency estimation (see methods in the main manuscript), with medium to large effect sizes. This was despite noisy latency estimates for the MTL (in particular, the ‘delta slope’ method could not be used to get meaningful Cohen’s d when comparing MTL to OCC). Latency estimates are also slightly higher for the ‘% of peak’ method than in the manuscript, as we estimated the latency at 25% of the peak (instead of 20% in the manuscript), again to allow meaningful estimations for the MTL.

      Author response image 1.

      In addition to this, we performed a simulation analysis where we statistically compared the measured PTL signals to ATL signals that have been artificially, incrementally, shifted forward in time. Author response image 2 shows (top row) the measured onset latencies differences between the 2 regions (PTL minus ATL) estimated using 4 different approaches as a function of the temporal shift applied to ATL, as well as the associated p-values (bottom row). As shown in Author response image 2, the ‘original’ unshifted data yields no significant difference between the 2 regions. The difference however becomes significant when ATL signals is shifted forward in time 10 to 30 ms, depending on the method used.

      Author response image 2.

      These 2 observations provide evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies, further supporting our claims.

      Last, as also suggested by reviewer 2, we conducted a thorough equivalence testing using ROPE and Bayesian factor to support the lack of differences between regions. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from early visual cortex than OCC). These analyses, now reported in the result section of the revised manuscript (Table 1), largely confirm the hypothesis of concurrent onset latencies across VOTC.

      (2) It is unclear why correlating mean timeseries helps understand how much variance is shared between regions (Figure 4). Any variance between images is lost when averaging time series across all images, and this metric thus overestimates the variance shared between areas. Moreover, the finding that correlating time domain signals across VOTC areas does not differ from correlating signals within an area could be driven by this averaging. For example, if the same analysis was done on electrodes in left and right V1 when half of the images had contrast in the left hemifield and the other half had contrast in the right hemifield, the average signals may correlate extremely well, while this correlation falls apart on a trial-by-trial basis. These analyses therefore need to be evaluated on a trial-by-trial basis.

      This is an interesting comment. We agree that variance between images is lost when averaging time-series across all images. However, to use the reviewer’s analogy, in order to support the claim that left and right V1 would show the same onset times and time-course (i.e., no hemispheric lateralization) for lateralized presentations (= the same kind of claim that we make in our paper), it’s the average response across images that should be compared, not a correlation run on a trial-by-trial basis (which would indeed falls apart because of a lack of response in the ipsilateral V1).

      Moreover, we would like to emphasize that the goal of this analysis in our paper is not to make claims about the variance shared between regions. In fact, this is not a key analysis in our paper, the outcome of which is not strictly necessary for the main argument made. Finally, if we were to perform a (time-consuming) image-by-image analysis in our study, correlations would be weak due to low signal-to-noise ratio (each face image appears only 1.6 times per stimulation sequence on average) and the fact that each face image appears after a different non-face image across presentations.

      (3) Previous studies on visual processing in VOTC have shown that evoked potentials are more predictive of the onset of visual stimuli than broadband activity (e.g. Miller et al., 2016, PLOS CB, https://doi.org/10.1371/journal.pcbi.1004660). Testing the prediction from a hierarchical representation that signals along the VOTC increase in onset time should therefore include an evaluation of evoked potential onsets in addition to broadband signals.

      We have used HFB responses in our study as these signals tend to be easier to characterize in the time domain than evoked potentials, and they are more local given their reduced SNR compared to evoked potentials (Jacques et al., 2022; https://pmc.ncbi.nlm.nih.gov/articles/PMC9457683/). Moreover we have previously shown highly correlated time courses across HFB and low frequency evoked potentials in the same paradigm (Jacques et al., 2022, eLife).

      Yet, to address this reviewer’s concern, we replicated the main analyses on low-frequency event-related potential signal, identifying contacts exhibiting significant face-selective responses in the same manner as in Jacques et al (2022). Namely, we start from bipolar-referenced sequences of recording corresponding to the full visual stimulation sequences (~70 s). For each recorded intracerebral contact, we average sequences in the time-domain, crop the average to contain an integer number of face frequency (1.2 Hz) cycles, run an FFT on the cropped sequences and identify the significant contacts with a Z-score procedure identical to that used for HFB signals. We then notch-filter out the visual response (6 Hz and harmonics) from the full length sequences, extract short epochs from the filtered sequences around the onset of each face image, average across epoch for each recording contact, subtract the mean signal measured in the baseline (-0.166 to 0 s relative to face onset) and take the absolute value (to be able to average across contacts despite differences of morphology and polarity). Significant contacts are then subjected to the same analyses as for the HFB signal.

      Results from these analyses are presented as supplementary material (Figure S9, Table S4) in the revised manuscript (referenced in lines of the main manuscript). While we were not able to obtain reliable latency estimates using the z-score method with the same parameters as for HFB signal, these analyses with ERP signal indicate similar onset latencies for ERPs as for HFB activity and largely replicate observations made with HFB. In particular, onset latencies were in a very similar range (~100 to 140 ms) with similar patterns across regions or along VOTC and between-region signal correlations. There were also a few significant face-selective responses over posterior ventro-medial occipital cortex, likely overlapping ‘early visual cortex’ (V1,V2v,V3v,hV4), probably due to limited low-level contributions in this paradigm (see Or et al., 2019, JOV; https://jov.arvojournals.org/article.aspx?articleid=2734585). Over these regions, onset latency was systematically earlier (up to 40ms) than in slightly more anterior regions, (i.e. anterior to -80 mm) where very little variability in onset latency was found up to the ATL region. We have acknowledged this in the revised manuscript.

      (4) Testing the second prediction, that the onset time of processing increases along the VOTC posterior to anterior path, is difficult using the iEEG broadband signal, because from a signal processing perspective, broadband signals are inherently temporally inaccurate, given that they are filtered. Any filtering in the signal introduces a certain level of temporal smoothing. The manuscript should clearly describe the level of temporal smoothing for the filter settings used.

      The reviewer is right that HFB signals are temporally smoothed, potentially yielding slightly underestimated onset latencies. However, our time-frequency analyses parameters ensured a minimal degree of smoothing. In fact, the original submission already contained a description of the expected temporal smoothing resulting from the wavelet transform. This is what we wrote in the original submission: “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had 20 ms of full width at half maximum (FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin), ensuring that onset timing information is accurate up to 10 ms (i.e half of the FWHM).”

      In the revised manuscript we further elaborate as follows:

      “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had a temporal spread of 20 ms (full width at half maximum - FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin). A simulation of HFB signals with a constant abrupt onset time and realistic signal-to-noise ratio indicates that the potential underestimation of onset latency due to the wavelet analysis is around 5-10 ms, which is on par with the value of the half width at half maximum (= FWHM/2 = 20/2 ms).”

      Author response image 3 displays simulated HFB signal (using identical wavelet parameters than in our manuscript) in an ideal scenario with a response starting at 150 ms in all trials (N=150 trials), reaching maximum 10 ms later. This provides a theoretical estimate of the slight underestimation of onset latency due to the wavelet transform. It shows onset latency estimates are at most 12 ms underestimation of true onset time.

      Author response image 3.

      That being said, given the physiological noise in the data, the fact that the response onset likely varies slightly from trial to trial, with a variable slope in activity increase, these wavelet parameters (within a certain margin) have likely little influence on the actual latency estimation.

      (5) The onsets of neural activity in VOTC are surprisingly early: around 80-100 ms. This is earlier than what has previously been reported. For example, the cited Quian Quiroga et al. (2023) found single neuron responses to have the earlier onset around 125 ms (their Figure 3). Similarly, the cited Jacques et al., 2016b and Kadipasaoglu et al., 2017 papers also observe broadband onsets in VOTC after 100 ms. Understanding the temporal smoothing in the broadband signal, as well as showing that typical evoked potentials have latencies compared to other work, would increase confidence that latencies are not underestimated due to factors in the analysis pipeline.

      In the revised manuscript, as suggested by reviewer 2, we have modified the data resampling strategy (using hierarchical bootstrap and permutation test that respects the nested structure of the data) to estimate onset latencies, confidence interval and permutation tests. Moreover, since the absolute onset latency estimates depend on the methods used, we now provide estimates using 4 different methods. The overall absolute onset latencies differ slightly across the 4 methods but all median onset latencies vary between 95 ms and 130 ms, which is similar to what was reported in some of the participants in Kadipasaoglu et al., 2017 (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0188834; note that latencies reported at the level of single sites or individual are usually higher due to lower signal-to-noise ratio). Moreover, these latencies for face-selective responses are actually very similar to those measured for non-selective/absolute responses to visual stimuli (e.g. Jacques et al., 2016: around 90-100 ms for faces in FG; Yoshor et al., 2007: ~100 ms in posterior Fusiform Gyrus; Regev et al. 2018: 90-120ms in posterior to middle FG). Other studies, measuring non-selective responses using more ‘conservative’ onset detection methods, find slightly later onset latencies (e.g. Cao et al. 2025: 139 ms in posterior FG; Martin et al., 2019: ~150 ms in ventral occipital -VO- regions).

      Cao R, Zhang J, Zheng J, Wang Y, Brunner P, Willie JT, Wang S. 2025. A neural computational framework for face processing in the human temporal lobe. Current Biology. DOI: https://doi.org/10.1016/j.cub.2025.02.063

      Martin AB, Yang X, Saalmann YB, Wang L, Shestyuk A, Lin JJ, Parvizi J, Knight RT, Kastner S. 2019. Temporal dynamics and response modulation across the human visual system in a spatial attention task: An ECoG study. Journal of Neuroscience 39:333–352. DOI: https://doi.org/10.1523/JNEUROSCI.1889-18.2018, PMID: 30459219

      Regev TI, Winawer J, Gerber EM, Knight RT, Deouell LY. 2018. Human posterior parietal cortex responds to visual stimuli as early as peristriate occipital cortex. European Journal of Neuroscience 48:3567–3582. DOI: https://doi.org/10.1111/ejn.14164, PMID: 30240547

      Yoshor D, Bosking WH, Ghose GM, Maunsell JHR. 2007. Receptive fields in human visual cortex mapped with surface electrodes. Cerebral Cortex 17:2293–2302. DOI: https://doi.org/10.1093/cercor/bhl138, PMID: 17172632

      As an important note, in the revised manuscript, we have removed data from 3 recording contacts from 1 participant that were located in the upper bank of the Calcarine Sulcus, which is actually outside of our VOTC region of interest. The 3 contacts being located in dorsal V1 or V2 were showing very early responses and were biasing our latency estimates for the OCC region.

      (6) Understanding the extent to which neural processing in the VOTC is hierarchical is essential for building models of vision that capture processing in the human brain, and the data provides novel insight into these processes.

      For additional context, a schematic figure of the hierarchical view and a more parallel system described in the paragraph on models of visual recognition (lines 553) would help the reader interpret and understand the implications of the paper.

      Our observations in the current study clearly indicates concurrent face-selective processing in the VOTC, which is incompatible with a serial hierarchical model. While we discuss how such concurrent activity could be implemented in the cortex (e.g., via direct input from ‘early visual cortex’ to different VOTC face-selective clusters), our data do not allow to provide more evidence in that respect to what already exists in the literature. Moreover, we are not providing data regarding connectivity (feedforward or re-entrant) either between face-selective regions or between these regions and ‘early visual cortex’.

      Author response image 4 shows a very simplified versions of standard hierarchical/serial versus concurrent/parallel models.

      Author response image 4.

      Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The data alone are valuable, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important, and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well-suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided. Collectively, the various analyses and rich illustrations provide superficially convincing evidence in favour of the conclusions.

      We thank the reviewer for their positive comments on our manuscript.

      Weaknesses:

      (1) The work was not pre-registered, and there is no sample size justification, whether for participants or trials/sequences. So a statistical reviewer should assess the sensitivity of the analyses to different approaches.

      Pre-registration of fundamental research in a clinical context is quite uncommon for intracranial data, owing, for instance to the time needed to accumulate data, or to the uncertainty of cortical sampling location in a given participant. Nevertheless, in the current study, sample size is much higher than in typical intracranial studies (usually 5-20 participants), in fact much higher than most typical Cognitive Neuroscience research. The same is true for the number of recording contacts (>10000 site here), and the number of trials considered for analysis. Each participant had a minimum of 164 face trials and an average of 262 trials (i.e. an average of 3.2 stimulation sequences of 82 trials), which is higher than most standard human electrophysiological studies.

      In the revised manuscript, to unsure that we have the maximum available power, and because our hypothesis is independent of hemisphere, we collapsed data across hemispheres for all analyses. We nevertheless provide analyses split by hemispheres as supplementary material.

      In addition, since onset latency estimations depend on the methods used, we now report onset latencies from 4 different methods (2 statistical and 2 non-statistical).

      (2) Frequentist NHST is used to claim lack of effects, which is inappropriate, see for instance:

      Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3

      Rouder, J. N., Morey, R. D., Verhagen, J., Province, J. M., & Wagenmakers, E.-J. (2016). Is There a Free Lunch in Inference? Topics in Cognitive Science, 8(3), 520-547. https://doi.org/10.1111/tops.12214

      Please see reply to the next comment (3).

      (3) In the frequentist realm, demonstrating similar effects between groups requires equivalence testing, with bounds (minimum effect sizes of interest) that should be pre-registered:

      Campbell, H., & Gustafson, P. (2024). The Bayes factor, HDI-ROPE, and frequentist equivalence tests can all be reverse engineered-Almost exactly-From one another: Reply to Linde et al. (2021). Psychological Methods, 29(3), 613-623. https://doi.org/10.1037/met0000507

      Riesthuis, P. (2024). Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science, 7(2), 25152459241240722. https://doi.org/10.1177/25152459241240722

      We thank the reviewer for pointing this out. In the revised manuscript we conduct and report a thorough examination of equivalence using ROPE and Bayesian factor to support the lack of differences between regions. We did not use TOST procedures as these have low power and require huge samples be meaningful (Riesthuis, 2024). Instead we relied on Bayes factor and descriptive proportion in ROPE. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from ‘early visual cortex’ than OCC). These analyses, now reported in the result section of the revised manuscript, along with effect sizes, confirm the hypothesis of concurrent onset latencies across VOTC.

      Riesthuis P. 2024. Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science 7.

      Detailed methods are reported as well:

      “In addition, to statistically assert whether onset latencies measured across main VOTC regions (OCC, PTL, ATL) were consistent with a concurrent (parallel) face-selective activation, we used Bayesian equivalence testing, relying on two separate metrics: (1) the percentage of differences in region of practical equivalence (ROPE), and (2) the Bayes factor using Cauchy prior. Equivalence bounds for ROPE were defined by combining two components: (1) a component of physiological variability and (2) a component reflecting expected delays in response onset between regions attributed to neural conduction delay, given the differential distances separating early visual cortex (EVC) from posterior face-selective regions (e.g. IOG) vs. anterior regions (ATL) and assuming signal mainly travels between VOTC regions through major postero-anterior axis fiber bundles of the Inferior longitudinal fasciculus (ILF) or the inferior fronto-occipital fasciculus (IFOF). Physiological variability corresponded to expected measurement noise and between-subject variability. Physiological equivalence bound was obtained by multiplying a Cohen’s d of 0.2 (conventional threshold for a negligible effect) with the (pooled) between-subjects variability in onset latency (computed for each region using a jackknife procedure and a correction factor of [N-participants – 1] to the jackknife standard deviation). For conduction delay bounds, the expected latency difference between regions under parallel activation depends on: (1) the distance between each region (considering the estimated origin of the signal is the same for all regions - EVC), and (2) neural conduction velocity. For inter-region distance we determined, for each region, the 5 and 95 percentile of the Talairach y-coordinate distribution and defined the maximal distance bounds as the distance between the y-coordinate corresponding to 5% of region 1 (e.g. OCC) to the coordinate corresponding 95% of region 2 (e.g. PTL). This resulted in the following maximum distance values: OCC-PTL: [54]mm; PTL-ATL: [54]mm; OCC-ATL: [83] mm. For conduction velocity, we used a constant value of 3.5 m/s, based on median axonal conduction velocity for cortico-cortical connections (Lemarachal et al., 2022; Van Blooijs et al., 2023), which was more conservative than using a range of values (e.g. 1.7 to 5.3 m/s based on Lemarachal et al., 2022). Maximum expected conduction delay was computed as [maximum distance / conduction speed] (e.g. for OCC to PTL: 0.054 / 3.5 = 15 ms). Resulting equivalence bounds (ROPE) were asymmetrical given than one region (e.g. OCC) is always closer to the source (EVC) than the other region (e.g. PTL) and were defined as [-1*physiologial_bound +1*physiologial_bound+max_conduction_delay]. For instance, using the z-score method to measured onset latencies, ROPE was [-12 to 27] ms for OCC to PTL, meaning that under equivalence, OCC can be activated up to 27 ms earlier than PTL (maximum conduction delay + noise), while allowing for some instances where OCC activates later (up to 12ms, due to noise only).

      For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (4) The lack of consideration for sample sizes, the lack of pre-registration, and the lack of a method to support the null (a cornerstone of this project to demonstrate equivalence onsets between areas), suggest that the work is exploratory. This is a strength: we need rich datasets to explore, test tools and generate new hypotheses. I strongly recommend embracing the exploration philosophy, and removing all inferential statistics: instead, provide even more detailed graphical representations (include onset distributions) and share the data immediately with all the pre-processing and analysis code.

      Data will be shared upon publication of the manuscript (see OSF repository in https://osf.io/2qzym). While we agree the dataset is large and could be explored in many ways, we do not consider the current study to be exploratory in nature. While our measurements could have turned out to clearly support hierarchical processing in human VOTC, our point in this manuscript is that the evidence derived from this large dataset unequivocally points instead toward concurrent activation of face-selective regions along the VOTC from IOG to antFG+ (i.e. a ~90 mm portion of cortex), with potential small variability accounted for by variability in axonal conduction velocity, signal-to-noise ratio or simple physiological variability. Other likely sources of variability such as type and density/size of fiber bundles across regions cannot easily be modeled with the current data set.

      (5) Even if the work was pre-registered, it would be very difficult to calculate p-values conditional on all the uncertainty around the number of participants, the number of contacts and the number of trials, as they are random variables, and sampling distributions of key inferences should be integrated over these unknown sources of variability. The difficulty of calculating/interpreting p-values that are conditional on so many pre-processing stages and sources of uncertainty is traditionally swept under the rug, but nevertheless well documented:

      Kruschke, J.K. (2013) Bayesian estimation supersedes the t test. J Exp Psychol Gen, 142, 573-603. https://pubmed.ncbi.nlm.nih.gov/22774788/

      Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin & Review, 14(5), 779-804. https://doi.org/10.3758/BF03194105 https://link.springer.com/article/10.3758/BF03194105

      All analyses and preprocessing stages are identical between regions and the number of trials is large enough not to be a constraining factor. As indicated above and below, we now report detailed equivalence testing and effect sizes and recomputed all statistics, taking into account the structure of the data as suggested by the reviewer.

      (6) Currently, there is no convincing evidence in the article to clearly support the main claims.

      Bootstrap confidence intervals were used to provide measures of uncertainty. However, the bootstrapping did not take the structure of the data into account, collapsing across important dependencies in that nested structure: participants > hemispheres > contacts > conditions > trials.

      Ignoring data dependencies and the uncertainty from trials could lead to a distorted CI. Sampling contacts with replacement is inappropriate because it breaks the structure of the data, mixing degrees of freedom across different levels of analysis. The key rule of the bootstrap is to follow the data acquisition process, and therefore, sampling participants with replacement should come first. In a hierarchical bootstrap, the process can be repeated at nested levels, so that for each resampled participant, then contacts are resampled (if treated as a random variable), then trials/sequences are resampled, keeping paired measurements together (hemispheres, and typically contacts in a standard EEG experiment with fixed montage). The same hierarchical resampling should be applied to all measurements and inferences to capture all sources of variability. Selectivity and timing should be quantified at each contact after resampling of trials/sequences before integrating across hemispheres and participants using appropriate and justified summary measures.

      The authors already recognise part of the problem, as they provide within-participant analyses. This is a very good step, inasmuch as it addresses the issue of mixing-up degrees of freedom across levels, but unfortunately these analyses are plagued with small sample sizes, making claims about the lack of differences even more problematic--classic lack of evidence == evidence of absence fallacy. In addition, there seem to be discrepancies between the mean and CI in some cases: 15 [-20, 20]; 8 [-24, 24].

      In light of the reviewer’s comment, we recomputed all timing analyses using a stratified hierarchical approach to evaluate confidence intervals (using bootstrapping), statistical comparisons (using permutation tests) and equivalence testing.

      This is what we wrote in the revised methods:

      “The first two timing parameters of face-selective response, onset and offset latencies, were quantified per main VOTC region using a hierarchical bootstrapping approach to respect the nested structure of the data (region > participants > contacts > trials). For each bootstrap iteration and each region, we first sampled participants with replacement. Within each sampled participant, we then sampled contacts and then trials within sampled contacts, with replacement. Resampled trials, then contacts within participants, then participants within a region, were successively averaged to obtain a bootstrapped region-level response from which we derived onset latency (4 different methods) and offset latency. We obtained bootstrap distributions of onsets/offsets using 2000 bootstrap iterations per region, allowing to compute the median and 95% confidence interval for these 2 parameters.”

      Then, later about permutation tests:

      “Statistical significance of latency differences between main VOTC regions was assessed using a hierarchical permutation test. We use a stratification approach to partition participants into a paired set (i.e. participants that had recording contacts in the two regions compared) and unpaired set (participants with contacts in a single region). For paired participants, the region labels were randomly swapped within subject (i.e., exchanging the signals from the two regions), thereby preserving all participant-, contact-, and trial-level structure while breaking the association between region and latency estimates. For unpaired participants, participants were randomly reassigned between regions while preserving the original group sizes, to generate pseudo-groups under the null hypothesis of no regional difference. In each permutation, signals were averaged across trials, then contacts, then participants within each permuted group and latency was computed and stored from the resulting region-level signals. We performed 10000 permutations to obtain a distribution of regional differences of latencies under the null hypothesis and determine the p-value as the fraction of the null distribution larger or smaller than the observed (non-permuted) difference.”

      And then about equivalence testing:

      “For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (7) Three other issues related to onsets:

      (a) FDR correction typically doesn't allow localisation claims, similarly to cluster inferences: Winkler, A. M., Taylor, P. A., Nichols, T. E., & Rorden, C. (2024). False Discovery Rate and Localizing Power (No. arXiv:2401.03554). arXiv. https://doi.org/10.48550/arXiv.2401.03554

      Rousselet, G. A. (2025). Using cluster-based permutation tests to estimate MEG/EEG onsets: How bad is it? European Journal of Neuroscience, 61(1), e16618. https://doi.org/10.1111/ejn.16618

      In fairness, we do not understand or share the reviewers’ concern here. Hundreds of fMRI or EEG studies use FDR or cluster tests to make inference about spatial or temporal location. We use FDR correction in one of the onset latency estimation method and only consider one-sided differences. Other methods in the revised manuscript do not use FDR correction.

      (b) Percentile bootstrap confidence intervals are inaccurate when applied to means. Alternatively, use a bootstrap-t method, or use the pb in conjunction with a robust measure of central tendency, such as a trimmed mean.

      Rousselet, G. A., Pernet, C. R., & Wilcox, R. R. (2021). The Percentile Bootstrap: A Primer With Step-by-Step Instructions in R. Advances in Methods and Practices in Psychological Science, 4(1), 2515245920911881.

      Again, we are not sure what the reviewer’s is referring to. The confidence intervals are computed on latency estimates from bootstrapped waveforms. In the revised manuscript, these waveforms are obtained by averaging (i.e. mean) resampled trials, resampled channels, resampled participants. A trimmed mean could not be applied in this condition, except perhaps when averaging across trials. But then the trimmed mean would have to be applied separately at each time sample which would disturbed within-, or between-trial, variability.

      (c) Defining onsets based on an arbitrary "at least 30 ms" rule is not recommended:

      Piai, V., Dahlslätt, K., & Maris, E. (2015). Statistically comparing EEG/MEG waveforms through successive significant univariate tests: How bad can it be? Psychophysiology, 52(3), 440-443. https://doi.org/10.1111/psyp.12335

      The rule of contiguous significant points is a heuristic that many researchers have used successfully to avoid spurious detection due to temporal autocorrelation. While we are aware that more sophisticated methods exist to correct for autocorrelation, such as cluster-based approaches, it is not directly usable since it requires comparing 2 conditions. The approach described in Piai et al., 2015 is interesting but incorrect as well since it relies on split-half simulations, which reduced signal-to-noise ratio, resulting in over estimated correction to be applied. In our revised manuscript, we rely on multiple methods to estimate onset latency, some of which not relying on this heuristic. Moreover, we apply plausible physiological constrains to our latency estimates, such as rejecting any onset before 40 ms after stimulus onset.

      (8) Figure 5 and matching analyses: There are much better tools than correlations to estimate connectivity and directionality. See for instance:

      Ince, R. A. A., Giordano, B. L., Kayser, C., Rousselet, G. A., Gross, J., & Schyns, P. G. (2017). A statistical framework for neuroimaging data analysis based on mutual information estimated via a Gaussian copula. Human Brain Mapping, 38(3), 1541-1573. https://doi.org/10.1002/hbm.23471

      (9) Pearson correlation is sensitive to other features of the data than an association, and is maximally sensitive to linear associations. Interpretation is difficult without seeing matching scatterplots and getting confirmation from alternative robust methods.

      We rely on Pearson correlation because this replicates the method used in Kadipasaoglu et al., 2017. It is also a widely accepted measure of (linear) relationship (in our situation we did expect linear or near linear relationships) in the literature. To address the reviewers concern, in the revised manuscript we nevertheless report, as supplementary material (Figure S10), the same functional connectivity analyses performed using the methodology and code provided in Ince et al. (2017). The results of this analyses are extremely similar to the results using Pearson’s coefficients.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 6, the response onset latencies are rendered in a smoothed manner on the brain surface. However, with this smoothing, variability between electrodes cannot be seen, and it would be better visualized in color, rendered in each electrode.

      Latency estimates computed at individual channels are noisy, which is why we do not report individual channels latencies but rather rely on averaging signals across contiguous channels, either across whole regions (Figure 4) or across smaller volumes as in Figure 6.

      (2) Onset latencies of 60 seconds seem extremely early compared to literature typically citing evoked responses with a latency of ~170ms. It would help if some additional sanity checks were shown, such as showing the latency of early visual responses. This would help with relative comparisons.

      In the revised manuscript the earliest median latency is 95 ms, which is in line with previous intracranial electrophysiology literature (e.g. Jacques et al., 2016; Jacques et al., 2022 ; https://pubmed.ncbi.nlm.nih.gov/26212070/; https://pubmed.ncbi.nlm.nih.gov/36074548/). The 99% confidence intervals can result in earlier latencies both due to some participants showing early responses and noise in latency estimates. Also please keep in mind that latency estimates are usually earlier when combining data across channels/participants compared to individual channels simply due to differences in SNR or across participants (see e.g. Kadipasaoglou et al., 2017).

      The reviewer indicates “…to literature typically citing evoked responses with a latency of ~170ms.”. We are assuming that they refer to the face-selective N170 ERP component measured on the scalp in EEG. Even with this ERP component, the face-selective response usually starts around 120-130 ms after stimulus onset (e.g. Rousselet et al., 2008; Jacques, Retter and Rossion, 2016; https://pubmed.ncbi.nlm.nih.gov/18831616/; https://pubmed.ncbi.nlm.nih.gov/27138205/) at scalp level. With the same highly sensitive paradigm as used here in EEG, we have systematically shown latency onsets of face-selective activity shortly after 100 ms (e.g., Retter et al., 2020; also Quek & Rossion, 2017) not accountable for by low-level visual cues (i.e., not present for phase-scrambled stimuli; Rossion et al., 2015; Or et al., 2019). Our latency onsets are also in line with spiking activity recorded with the same approach in the LatFG (Laurent et al., 2026) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5955677

      (3) Figure 6B shows that the variability of latencies in the ATL is larger than the variability of latencies in the PTL. It would be helpful to evaluate whether, rather than in mean onset latencies, there is a change in variability in onset latency along the VOTC.

      This is an interesting point. However, it is difficult to evaluate since SNR is reduced in the ATL compared to OCC or PTL (Jacques et al., 2022). As a result, any measured modulations in the variability of onset latencies along VOTC may simply reflect changes in the precision of latency estimation driven by SNR variability.

      (4) Line 415 typo: 'there appears to be no delay' instead of 'there appear to be no delay'.

      We thank the reviewer for their careful reading of our manuscript. This has been corrected.

      (5) The discussion states in lines 455-457 that "a large proportion of neuronal populations in anterior VOTC regions exhibiting similar activity to different face images independently of the context in which they appear". However, this claim about similar activity to different faces should be evaluated and tested at the single-trial level. In addition, it is not clear how context was varied in the experimental design.

      Context is variable because each face image appears directly after a different object image (or object images) in the sequence. We have shown also in previous studies with this paradigm in EEG that the time-course of face-selective responses is similar across base frequencies (3-15 Hz) unless the rate is too fast, and whether an orthogonal or explicit face categorization task is used (Retter et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32119982/). Note that we do not claim that activity is identical across images but similar – if it was not (largely) similar, the averaged response would be jittered and low.

      (6) Line 516-517: DTI does not provide evidence for whether connectivity is direct or not, and what the directionality of connectivity is between two areas. This sentence should therefore state "..., suggest independent connections between early visual cortex and face-selective regions...".

      This has been rephrased.

      (7) Line 555 in the discussion, the definition of low-level visual should be expanded to include other early visual areas that have been demonstrated to respond earlier than VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018), to avoid the suggestion that V1 directly projects synaptically to all of VOTC (e.g. Markov et al., 2014, Cerebral Cortex, https://doi.org/10.1093/cercor/bhs270).

      We are not proposing that V1 directly projects directly/synaptically to all of VOTC, i.e., without other low-level retinotoptic areas involved; only that face-selectivity in the association cortex is not organized hierarchically. We have revised this sentence.

      (8) Line 585, for the sentence: "with temporal synchrony strengthening their connections", evidence or citations should be provided.

      Citations have been provided.

      (9) It is not clear what is meant in the paragraph starting in line 571: do the authors suggest that top-down signals are not necessary for fast recognition of clear views of faces, or additionally argue that these top-down signals are not necessary for detecting ambiguous or degraded inputs as faces?

      Exactly: That top-down (i.e., descending) signals may contribute but would not be necessary for fast recognition of clear views of faces AND for detecting ambiguous or degraded inputs as faces.

      (10) No statement was provided on data or code availability.

      Data will be made available on a repository upon publication (see https://osf.io/2qzym).

      Reviewer #2 (Recommendations for the authors):

      (1) FDR correction: which one? Please provide a reference.

      We now provide a reference, both in the results and methods: Benjamini and Hochberg, 1995.

      (2) In the introduction, this statement is too strong: "arguably the most familiar and ecologically valid stimulus". It is unclear how static 2D representations of faces are the most familiar and valid stimuli. Could you rephrase this? What about other very familiar stimuli like letters, words and biological motion?

      This statement is not about static 2D images of faces, but faces in general (in their natural environment). We do consider human faces (in general, not restricted to laboratory context) to be indeed the most familiar and ecologically important stimulus, both from an ontogenetic and phylogenetic perspective, unlike written material.

      (3) About the questioning of a strict temporal hierarchy, this EEG reference comes to mind: Foxe, J. J., & Simpson, G. V. (2002). Flow of activation from V1 to the frontal cortex in humans. Experimental Brain Research, 142(1), 139-150. https://doi.org/10.1007/s00221-001-0906-7

      As confirmed by Foxe et al. ’s (2002) paper to which the reviewer is referring to, there is indeed ample evidence that areas in the dorsal stream or frontal cortex (e.g. FEF) are activated very soon after V1 and before many ventral stream regions (e.g. Lamme and Roelfsema, 2000; https://pubmed.ncbi.nlm.nih.gov/11074267/). While Foxe et al.’s 2002 is highly valuable, it can hardly be compared with our current study which looks specifically into ventral stream areas which are largely indistinguishable using scalp EEG as in Foxe et al. ’s paper.

      (4) Regarding statistical significance, there is no such thing as a "trend". The threshold for a trend should have been pre-registered and applied to both sides of the magical boundary, for instance, with matching conclusions for a "trend toward non-significance (p=0.04)". P values near 0.05 provide weak support against the null. I would suggest leaving it at that. Nothing special happens at 0.05.

      This no longer appears in the revised manuscript.

    1. eLife Assessment

      This important study provides convincing evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial. By combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia.

      Comments on revised version:

      Thank you for the opportunity to re-review this interesting paper.

      This reviewer struggled to follow along with the author's response letter. It was challenging because responses were brief and did not list the specific changes made, leaving the reviewer to search for changes in the text. As best I could tell, there were no changes made that matched some of the highest-priority suggestions.

      Critique #1 - This reviewer did not find the presentation of confidence intervals in the Abstract and other sections, which were suggested. Please note the format that was suggested in the original critique from Reviewer 1.

      Critique #2 - regarding removal of unsubstantiated claims and use of a p-value to compare analysis to a null hypothesis - it seems the authors skipped over this critique and did not address it.

      Regarding bias: the authors answered a different question than the one asked. The reviewer asked what proportion of transmissions were sampled; the authors stated that only communities from which phylo data was acquired were modeled. Was sampling 100% in those communities? Please provide the percentage and provide analysis that shed light on how sampling bias could impact the analysis.

      Regarding "cherries" - the reviewer did not understand the author's response. The query was regarding what percent of the total number of phylogenetic pairs (denominator) were the 355 that had high confidence in directionality (numerator). The response could be expressed be a proportion.

      The expectation of ART reducing the age of sources of transmission seems unrealistic to this reviewer. People on ART are not always adherent and can still transmit during gaps in adherence. ART dramatically increases life expectancy with HIV, which would have the opposite effect.

    3. Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex specific transmission dynamics in Zambia with a goal of allocation of resources.

      Strengths:

      Important analysis to hone in on key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches used, and while the phylogenetic data was markedly more limited, it mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses and providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce.

      Comments on revised version.

      The revised manuscript clarifies the impact and utility of this work and better allows the comparability of the two methods. Highlighting the differences (or lack thereof) between the undiagnosed and diagnosed population) simplifies the public health approach.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial; by combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength. However, the evidence supporting key conclusions is incomplete, with several claims insufficiently substantiated by the data presented. Improvements in data presentation (e.g., quantification of qualitative statements, statistical estimates, and clearer description of results) would substantially strengthen the paper.

      We thank the editor and reviewers for their positive comments. We have revised the manuscript in response to the points raised, as described below.

      First of all, we would like to summarise what we have changed regarding the presentation of summary statistics throughout the text. We agree that many of the statements in the original submission tended towards being qualitative. This was the result of shying away from presenting two separate estimates, with different ways of quantifying uncertainty, in the text. The phylogenetics could be presented as mean and confidence interval, while the IBM would need some measure of centrality (mean or median) and the highest density interval for a summary statistic (e.g. the mean age gap) as it varies over the posterior. These are not directly comparable. We have now changed this to present both where appropriate, with cautionary note about the difference between the CIs and HDIs (lines 257-260).

      We also were somewhat arbitrary regarding where we chose to summarise the posterior in the IBM or look at the best-fitting single simulation, and where we presented the mean as opposed to the median. We have done a considerable overhaul of what is presented in this revision:

      (1) We always present the posterior summary unless the level of detail is such that summarising uncertainty over the posterior is not feasible (e.g. in figures 3, 4 and 5). In the latter case we still use the best-fitting IBM replicate.

      (2) In the main text we always present the mean. For the phylogenetics the summary statistics are mean and confidence interval. For the IBM this is the posterior mean, and 95% HDI, of the mean of a particular statistic as calculated in each of the 1000 IBM replicates. For example, each replicate will have its own distribution of male source ages which have a mean value. These means also vary over the posterior, and a mean of them is calculated, as well as the HDI interval to represent posterior uncertainty. This “mean of means” may be a slightly confusing piece of terminology at first glance, but it allows us to properly capture posterior uncertainty in a way we mostly avoided in the first submission.

      One result of 1) above is a change to figure 6. It is now summarised over the posterior, with the result that time trends that were previously not evident become clear. This changes our conclusions slightly (lines 529-537) but it should be noted that the magnitudes of the trends remain small.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia. The current manuscript draft is methodologically strong, but needs revision to strengthen the take-home messages. As written, there are many possible take-away conclusions. For example, the agreement between IBM and phylogenetic analysis is noteworthy and provides a methodological focus. The revealed age patterns of transmission could be a focus. The effects of the PopART intervention and the consequences of a 1-year disruption could be a focus. It is important, though, that any main messages summarized by the authors are substantiated by the evidence provided and do not extrapolate beyond the data that have been generated. I recommend that the authors think deeply about what the most important, well-supported messages are and reframe the discussion and abstract accordingly.

      We have rewritten the abstract, and also made changes to the discussion in order to centre our message around the contribution of particular of demographic groups to transmission, and how, with that contribution revealed, such groups can be selected for specialised interventions.

      Strengths/weaknesses by section:

      (1) ABSTRACT

      The Abstract summarizes qualitative findings nicely, but the authors should incorporate quantitative results for all of the qualitative findings statements.

      The abstract in the revision is extensively revised, and contains quantitative estimates throughout, from both methodologies where appropriate.

      The ending claim is not substantiated by the modeling scenarios that have been run: "targeted interventions for demographic groups such as under-35 men may be the key to finally ending HIV." It is straightforward to run this specific scenario in the model to determine whether or not this is true.

      Our modelling framework is not set up to model the “last mile” of HIV elimination, notably as it has no component for MSM or FSW transmission, and we do not feel that we could confidently present results regarding it. As a result, this statement has been greatly softened in the new abstract (lines 75-78).

      The authors should add confidence intervals to the quantitative metrics, such as the 93.8% and 62.1% incidence reduction.

      These have been added.

      (2) RESULTS

      The authors should check the Results section for any qualitative claims not substantiated by the analyses performed, and ensure the corresponding analyses are presented to support the claims.

      The Results and Methods describe the model's implementation of the PopART intervention differently. The Methods describes it as including VMMC, TB, and STI services, while the Results only mentions intensified HIV testing and linkage.

      This is a slight misreading of the text. That paragraph in the Methods is describing the trial itself, not the modelling framework.

      A limitation of the model is that HIV disease progression is based on the ATHENA cohort in the Netherlands, which is a different HIV subtype (B) than the one in the research setting (C). The model should be configured using subtype C progression data, which have been published, or at least a sensitivity analysis should be conducted with respect to disease progression assumptions.

      The available literature does not suggest a significant difference in progression between subtypes B and C, and we have added text and citations to this effect (lines 699-701).

      In Table 2, the authors should consider adding a p-value to establish whether or not IBM and phylogenetics estimates are different.

      We have done this; the appropriate test was a posterior predictive check. See lines 261-263, 575-579 and 805-814.

      (3) DISCUSSION

      The literature review and comparison of study results to previously published phylogenetic studies is very nice. The authors could strengthen this by providing quantitative estimates with CIs for a more scientific comparison of the study results vs. prior studies, perhaps as a table or figure.

      We have expanded the discussion on this point (lines 504-527). We considered adding a table, but the existing literature that directly answers the questions we ask is quite limited and fragmentary. For example, Monod et al do not present a complete treatment of age gaps. The literature using regression analyses to identify predictors of HIV prevalence or incidence related to partner age is extensive, but those results are not directly comparable to ours.

      The authors state that due to "the narrow geographical catchment area... The results should not be automatically extrapolated to apply to other SSA settings." The authors should exercise this caution when comparing the results to studies in South Africa and elsewhere.

      We have made more explicit acknowledgements of these limitations (lines 598-600).

      There are many other limitations to the analysis, including some mentioned above, that are not acknowledged. The authors should think carefully about what the most important limitations are and acknowledge them honestly at the end of the Discussion section.

      The limitations paragraph has been revised (lines 598-605).

      Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex-specific heterosexual HIV transmission dynamics in Zambia, with the goal of allocating resources.

      Strengths:

      Important analysis to hone in on the key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches were used, and while the phylogenetic data were markedly more limited, they mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses. The authors did a nice job of providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce

      Weaknesses:

      To increase the impact and utility of this work, it would be helpful to parse the analysis just a bit further to estimate the roles of undiagnosed vs diagnosed and untreated subpopulations on this transmission. PopART is a multifaceted intervention, but the cost, effort, and approach to reengagement in care vs testing/treatment can be quite different.

      We have now provided stratified results by diagnosed and non-diagnosed status of the source, as well as an overall summary of the proportion of undiagnosed sources by age and sex. See lines 305-310, 539-547, and table 3.

      Recommendations for the authors:

      Reviewing Editor:

      We commend you for conducting a rigorous and comprehensive study titled "The age and sex dynamics of heterosexual HIV transmission in Zambia: an HPTN 071 (PopART) phylogenetic and modelling study" that significantly advances the understanding of HIV transmission dynamics in sub-Saharan Africa. The study utilizes an innovative dual-methodology approach integrating individual-based mathematical modelling (IBM) and pathogen phylogenetics to characterize heterosexual HIV transmission patterns by age and sex during the PopART trial in Zambia.

      This manuscript reports on HIV transmission dynamics in Zambia using data from the PopART study, combining individual-based modelling and phylogenetic analysis. The use of two independent methodologies enhances confidence in the consistency of the findings and enables robust cross-validation. The work addresses an important topic in HIV prevention, particularly in settings where resources may become more constrained, and offers insight into potential demographic targets for intervention.

      However, several aspects of the manuscript limit its current impact. The main take-home messages are diffuse and not clearly presented. Some conclusions in the abstract and discussion appear to go beyond the scope of the presented data. For instance, the claim that targeting under-35 men may be key to ending HIV is not directly tested in the modelling scenarios and should be reframed or removed unless supported by new analyses. Furthermore, important quantitative details, such as confidence intervals, p-values, and precise age group estimates, are lacking in key sections (e.g., the Abstract and Results).

      The authors are encouraged to clearly identify and communicate their central findings, ensure all claims are fully supported by their analyses, and make the data more accessible to readers by adding detailed, quantitative summaries where needed.

      The following are our recommendations to the Authors:

      (1) Clarify Study Objectives and Central Messages

      Reframe the abstract and discussion to highlight a clear, well-supported set of main findings.

      Avoid overgeneralized or unsubstantiated claims, especially those not directly tested by your model (e.g., the effectiveness of targeting under-35 men).

      As stated above, we have revised this text accordingly.

      (2) Support Qualitative Claims with Quantitative Data

      Provide numerical results, including effect sizes and confidence intervals, wherever qualitative trends are mentioned.

      For example, restate: "The largest gaps for female recipients were among the youngest" as "... in the age group XX-YY with OR = Z.Z (95% CI: A.A-B. B)."

      As mentioned at the top of the review, we have overhauled the treatment of summary statistics extensively, and now give confidence or highest density intervals throughout the text.

      (3) Improve the Results Section

      Check that all claims are supported by the analyses, and ensure figure references are accurate.

      The statements that went beyond what was supported, notably about ending the epidemic by targeting young men, have been removed. The typo in table references has been fixed.

      Annotate Figure 6 with trendline coefficients and p-values where applicable.

      The takeaway message of figure 6 has now changed and we no longer see no trend, just a minor one.

      Revise Figure 4 for clarity or consider replacing it with a tabular format.

      We would prefer to keep the current figure 4, as we have not found any clearer way to illustrate the patterns, which are the consequence of the phenomenon observed in figure 5. We have put more explicit descriptive text in the discussion, linking the two figures (lines 470-476).

      (4) Address Potential Bias and Model Assumptions More Rigorously

      Explain sampling bias in IBM and phylogenetics (e.g., how the 355 high-confidence phylogenetic pairs were selected).

      The reviewer comment regarding the 355 pairs was based on a misapprehension; we used all the pairs we found using the phyloscanner pipeline. There are no sampling bias issues involved in the IBM as every individual in the simulations is considered. Appendix 2 includes some sensitivity analysis results if the procedure used to find the 355 is changed.

      Discuss how the use of subtype B disease progression data from the ATHENA cohort may impact results in a subtype C setting. A sensitivity analysis would strengthen this.

      Subtype B progression data was used in the absence of any appropriate data from subtype C, but the literature does not suggest any major difference between the two (lines 699-701).

      (5) Include More Detail on Undiagnosed Populations and ART Effects

      Estimate the roles of undiagnosed and untreated subpopulations in driving transmission.

      As mentioned above, this analysis has been added.

      Clarify mechanistically how ART might influence age gaps in transmission dynamics.

      This now is clarified in the introduction (lines 127-129).

      (6) General Improvements

      Provide p-values where comparisons are made (e.g., in Table 2).

      Use consistent terminology and definitions across Methods and Results.

      Add more discussion on limitations, especially regarding generalizability to other SSA settings.

      All of these have been inserted as previously mentioned.

      By addressing these points, the manuscript would present a more coherent narrative and a stronger, evidence-based contribution to the field. We appreciate you all for your fantastic effort and hope you will reflect the feedback in your final paper.

      Reviewer #1 (Recommendations for the authors):

      Thank you for the opportunity to review this interesting manuscript.

      In the public review, I have recommended that the authors should incorporate quantitative results for all of the qualitative findings statements. As one example, I would recommend that "We found the largest gaps for female recipients were among the youngest of those recipients" is re-written as "The largest gaps for female recipients were in the age group XXX-YYY with OR=ZZZ (XXX-YYY)." such as odds ratios, and specific outcome definitions including ages. To give one more example: "immediate increase in the average age at transmission of both sources and recipients" could be rephrased as "increase in the average age at transmission by XXX (YYY-ZZZ) years for sources and XXX (YYY-ZZZ) for recipients over [TIME PERIOD]."

      We hope the revisions we have made to the statistical presentation are satisfactory as a response to this request.

      Again in the public review, I recommended checking the Results section for any qualitative claims not substantiated by the analyses performed, and ensuring the corresponding analyses are presented to support the claims. An example is: "Trends are minor or non-existent in the former two variables." - please annotate Figure 6 (assuming the authors meant to reference Figure 6 and not 7 here?) to show over what period trendlines were fit and provide the coefficient and CI. To support the stated claim even more strongly, a p-value might be apt with a null hypothesis of a slope of zero.

      Please check the numbering on all figure references in the text, as some appear to be misnumbered. E.g., where the text refers to Figure 7, I believe the authors meant to reference Figure 6.

      The change to how we handled the statistics has changed the message of figure 6 (which is now figure 7) and rendered this somewhat moot. We have checked that all figure and table references are now correct.

      Figure 3 is very nice, but if the axes were flipped on one panel, it would make them easier to compare, and then adding some statistics to assess whether the patterns are the same or different when a man vs woman is the source.

      We have flipped the axes here.

      Figure 4 was too complicated for me. I could not follow the Sankey flows because there is too much going on and overlapping. Consider revising to make it easier to digest... perhaps to table format?

      As mentioned above, we would prefer to keep this figure, but we have situated it better in the text.

      Reviewer #2 (Recommendations for the authors):

      A few points that would improve the clarity and the strength of the manuscript

      (1) There is a need to clarify more about how the IBM and phylogenetic data does not suffer from sampling bias. For e.g.,

      Line 205: What proportion of the transmissions modeled in the IBM from Zambia?

      All of them. We confined the analysis of the IBM to the Zambian communities from which phylogenetic data was acquired (lines 755-758).

      Line 217: What proportion of the phylogenetic pairs (cherries) suggesting transmission were the 355 that had high confidence in directionality. How do these pairs compare to the others

      There was no identification of “cherries” involved in picking these pairs; the phyloscanner procedure does not use that step. We confined our analysis solely to the pairs for which we did identify a direction of transmission; that is the 355. Appendix 2 includes a sensitivity analysis involving varying the parameters by which these were identified.

      (2) I appreciate the authors noting that MSM transmissions are unlikely to be playing a role in this cohort, as noted in previous work by the group. However, systematic undersampling of men is common in other study cohorts of HIV. While the MSM and heterosexual networks may be relatively distinct, undersampled men who are bridging the networks could impact the estimates. Can the authors use the time to diagnosis analysis (HIV phyloTSI) to estimate rates of undiagnosed men and women?

      We feel that this is beyond the scope of this work. The phylogenetics dataset in its totality could be used for this purpose (although it is probably highly biased towards undiagnosed individuals due to the considerable majority of samples coming from the healthcare facilities). However, we concentrate here solely on the subset involved in our probable transmission pairs, which is fairly small. Extending the scope to an exploration of the full dataset would seem like a separate study, which we do have plans to do.

      We have used the IBM for this question instead (lines 303-321), however, as MSM transmission was not modelled, it is also not ideal for answering this question. Ultimately we feel that the way these studies were implemented makes it an unsatisfactory tool for answering the MSM question, important as it is.

      (3) Expanding on the point above, in other settings, transmission to young men has been associated with partnerships with older men, and if these young men then transmitted to young women, would we see a similar effect as noted in these models (assuming the young men were less well sampled).

      Our previous work (Hall et al., 2024) suggested no excess of identified male-male pairs in the phylogenetics dataset which might suggest cryptic male-to-male transmission. The age disparities would be worth exploring had this been found, but is curtailed by the lack of it.

      (4) Related to the point above, is there an estimate of the populations (age and sex) that are undiagnosed in the IBM model? Can this be teased out... is transmission from men to women more likely 2/2 lack of diagnosis... or lack of engagement in care?

      We have explored results by diagnostic status as it pertains to age and sex, but we feel that moving on to a more general exploration of the role of diagnosis and lack of engagement in care is again going beyond the scope of what is already a long paper.

      (5) I'm still not fully clear as to why ART might affect age gaps. Can this be explained in more detail?

      See lines 127-129.

    1. eLife Assessment

      This valuable study describes PXGS, a poly-transgene expression system that exploits the mutually exclusive splicing of Dscam variable exon 4 to enable conditional, simultaneous expression of up to 12 transgenes in Drosophila, addressing a longstanding limitation in which conditional co-expression has been restricted to a handful of genes. The approach is conceptually elegant and technically accessible, with potential applications spanning neuroscience, synthetic biology, and biomanufacturing across arthropod species. The evidence that Dscam exon 4 splicing is preserved in a UAS vector and that individual alternates can be replaced with functional transgenes is solid, and the in vivo axonal re-wiring application provides a convincing proof of principle. Quantitative characterization of expression levels, a direct demonstration of expression across all twelve positions, and additional imaging controls would further substantiate the system's utility and scope.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes the development of an expression system enabling up to 12 transgenes using the alternatively spliced fourth exon of *Drosophila* *Dscam* gene under the control of a UAS element. This will be a useful tool if expression is needed in *Drosophila* cells (in culture or in vivo). where *Dscam* splicing machinery is active, which limits its use.

      Strengths:

      The tool developed is based on a well-established genomic element. The underlying idea is relatively simple yet effective.

      Weaknesses:

      The authors describe the weaknesses of their system well, most importantly, depending on the presence of adequate levels of Dscam splicing factors in targeted cells. This likely limits effective use of the methodology to some cell lines (e.g., S2) and certain tissues (nervous system and innate immune system). The manuscript could do a better job in showing protein expression levels more quantitatively, either in comparison to other methods or as absolute values (transcript numbers, protein molarity, etc.).

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Yu et al seek to develop a Drosophila genetic tool to simultaneously co-express up to 12 transgenes. They leverage the native Dscam exon 4 alternative splicing to generate a UAS to enable cell- and temporal-specific expression of transgenes. This tool is called the poly-transgene expression system (PXGS). Previous approaches to co-express transgenes have been limited to four to five genes, so PXGS would be a significant advancement, especially when examining processes that require robust expression of many genes to confer function. The authors showed that PXGS can drive expression of multiple (1) fluorescent reporters and (2) cell surface receptors in different cell types (neurons, glia, and muscles). However, there are major proof-of-principle experiments missing to demonstrate the utility of PXGS and its potential limitations. Additionally, some of the data is just not interpretable, and experimental rigor is significantly lacking.

      Strengths:

      Developing a genetic tool to co-express transgenes beyond what is currently available would be significant.

      Weaknesses:

      (1) While the authors stated that each PXGS construct can express 12 transgenes, this was not directly tested - the largest number of genes tested was in the PXGS_fluorophores, which has 4 genes inserted in 10 alternates (and therefore it can be determined if genes from all the alternates are spliced in at a meaningful level).

      a. First, the authors state that they tested the expression of the fluorophores in S2 cells using RT-PCR before generating the fly line. However, this data is not shown. Also, it is possible to test expression and localization of UAS transgenes in S2 cells with a ubiquitous GAL4, similar to what they did for GFP expression in Supplemental Figure 1.

      b. In the corresponding figure for this experiment (Figure 2), there are some concerning expression patterns and potential channel bleed-through/crosstalk. The nSyb-Gal4 is a pan-neuronal driver, yet expression of three fluorophores was extremely minimal. This could potentially be explained by the deterministic vs random alternative splicing. However, it is more concerning that in the GFP, RFP, and iRFP channels, the exact same tiny cluster of neurons is observed, suggesting potential bleed-through of the channels. Appropriate controls are required, including expression of a traditional fluorescent reporter with the nSyb-Gal4 they are using. And replicates would help with the experimental rigor.

      c. Additionally, it would be helpful to know if genes inserted at each alternate exon are expressed at a similar efficiency (vs. some alternate exons have higher levels of expression)

      (2) PXGS expression in non-neuronal cells: The authors attempt to show that fluorophores targeted to different cellular compartments can be expressed in neurons and non-neuronal cells (glia and ubiquitously).

      a. In Supplemental Figure 2A, they first use S2 cells to confirm expression, which does confirm. However, they use no markers to show that the fluorophores localize to the corresponding compartments (e.g., mitochondria and nucleus). In Supplementary Figure 2B, with that magnification and resolution, it is impossible to determine if the fluorophores localize properly.

      b. In Figure 3, it is impossible to know if there is any glial expression based on those images. They state that "subcellular localization of fluorophores was observed in the flight muscle", but again, that cannot be concluded from the images. Also, the schematic of the construct is the same one used in Supplementary Figure 2. Why show it again?

      (3) In the functional expression of PXGS transgenes section, while the authors used RT-PCR to show that each receptor gene is transcribed in S2 cells, it was not tested if they are correctly expressed, translated, and localized in the fly. The functional outcome observed (Figure 4 b-c) could be the result of misexpression of one or multiple genes.

      a. Supp Figure 3: Why does the Sli lane have so many bands?

      b. Why were these genes chosen for misexpression? Is there any evidence that they are required for the wiring of the mechanosensory neuron? Does the co-expression lead to an additive effect?

      c. The RT-PCR result (Supplemental Figure 3) showed variations (e.g., kek and kir being significantly dimmer than tutl; multiple products for Sli). Is there any explanation behind this, and could this be the outcome of some alternate exons being more efficiently spliced than others?

      d. Figure 4: These images seem to be taken with a widefield scope and only one plane. Is it possible that some of the pSC neurons are in a different Z plane, and they are not being captured here? There definitely is part of the axon terminal out of focus in some of the images. Also, most of the figure graph axes (e.g., 4b) are extremely difficult to read. And the figure overall is not easy to interpret.

      (4) Supplemental Figure 5: This figure is quickly mentioned in the Discussion without much explanation. First, this must be in the Results section since it is an experiment. Second, this needs more context because, as is, it seems like it was just thrown into the manuscript.

      (5) The authors mentioned that the size of the inserted genes could be a limitation for this technique and tested cell surface receptors of different sizes. However, there was no explicit discussion in the main text.

    4. Reviewer #3 (Public review):

      This paper, by Brian Chen and collaborators, adapts the highly alternatively spliced Dscam1 gene locus for use in a system for simultaneous multi-transgene expression in a variety of insect species. Specifically, they show that the hypervariable Dscam1 exon 4 region maintains its alternative splicing when placed in a UAS expression vector, and that each of the twelve exon 4 alternates can be replaced with an exogenous gene such that co-expression of up to twelve proteins can be achieved. Since the co-expression of more than a few proteins simultaneously is difficult, this represents a significant advance with multiple use cases. The authors validated the technique by assessing expression in vitro and in vivo, and by rewiring Drosophila sensory neuron axons by simultaneously expressing several cell surface receptors within the neuron. Overall, this is a clearly written paper that describes a potentially important new system. I have no major criticisms.

    5. Reviewer #4 (Public review):

      From the Reviewing Editor:

      All three reviewers recognized the conceptual originality of PXGS and its value to the Drosophila community and the broader multi-gene expression field. The core demonstration - that Dscam exon 4 mutually exclusive splicing is maintained in an exogenous UAS vector, and that individual exon alternates can be replaced with genes of interest for conditional in vivo expression - was viewed as solid and creative. The in vivo application re-wiring pSc axonal arbors using PXGS constructs loaded with cell surface receptors was noted as an encouraging functional validation.

      The reviewers differed in their overall enthusiasm. Reviewer 3 found the experiments straightforward, the results clear, and the system a significant advance with broad use cases, with no major criticisms. Reviewer 1 viewed the evidence as broadly solid, with the principal limitation being a lack of quantitative expression data and a dependence on adequate Dscam splicing-factor levels that constrain the system's applicable cell types - a limitation the authors themselves describe well. Reviewer 2 was the most critical, finding the underlying concept significant but the supporting data insufficiently rigorous to conclusively establish the tool's utility. The points below reflect the areas where reviewers - principally Reviewers 1 and 2 - felt the manuscript could be strengthened, should the authors choose to revise.

      (1) Quantitative characterization of expression. Reviewers 1 and 2 both noted the absence of a quantitative comparison of PXGS-driven expression - in absolute terms (transcript numbers, protein amounts) or relative to standard UAS constructs. Given that signal is inherently divided across 12 alternates per transcription event, characterizing expression efficiency and whether all alternates are spliced and expressed at comparable levels would substantially strengthen the manuscript. Reviewer 2 specifically asked whether genes at different exon 4 positions are expressed with similar efficiency.

      (2) Direct demonstration of 12-transgene expression. Reviewer 2 noted that, although the manuscript claims expression of up to 12 transgenes, this was not directly tested - the largest construct placed 4 distinct genes across 10 alternates, and the largest functional test used 3 genes per construct. Either a direct demonstration with more positions occupied or a more carefully bounded claim supported by the probabilistic framework would address this.

      (3) Controls and interpretability of fluorophore expression. Reviewer 2 raised concerns about Figure 2, where expression of three fluorophores under nSyb-Gal4 was minimal, and the same small neuronal cluster appeared across the GFP, RFP, and iRFP channels - raising the possibility of channel bleed-through. Appropriate controls (including a conventional single UAS-fluorophore driven by the same nSyb-Gal4) and replicates were requested. Reviewer 2 also noted that the S2 cell validation data for the fluorophore constructs, described as having been performed prior to fly line generation, are not shown.

      (4) Non-neuronal expression and subcellular localization. Reviewer 2 noted that the compartment-specific localization claims (mitochondria, nucleus) in Supplemental Figure 2 and Figure 3 are not supported by co-markers, and that the magnification and resolution in key panels are insufficient to confirm proper localization, particularly in glia.

      (5) Functional expression of receptor constructs. Reviewer 2 noted that, while RT-PCR confirms transcription of each receptor in S2 cells, correct translation and localization in the fly were not directly tested, so the observed phenotypes could reflect mis-expression of one or a subset of the genes. Clarification of why the specific genes were chosen, whether co-expression produces additive effects, and the cause of the variable RT-PCR band patterns (e.g., the multiple Sli products, dimmer kek and kir signals) was requested.

      (6) Supplemental Figure 5 (synthetic biology / RNAi). Reviewer 2 noted that this figure is mentioned only briefly in the Discussion despite representing an experiment, and recommended moving it to the Results with appropriate context. Reviewer 1 separately queried the meaning of the two white boxes in this figure and whether the RT-PCR convincingly supports expression of all genes shown.

      (7) Figure quality and labeling. Both reviewers flagged that the labels in Figure 4 (particularly panel d) and the axes in Figure 4b are too small to read. Additional labeling points were noted (alignment of "Repo-GAL4" and "brain" in Figure 3; unlabeled images in Supplemental Figures 2 and 4; clarification of whether the two lanes per group in Figure 1c are replicates or use different primers). Reviewer 1 also suggested that some figure legends describe conclusions rather than what is shown, and recommended that legends describe the data with interpretation kept to the text.

      (8) Additional points. Reviewer 1 suggested showing more of the gel in Figure 1 to demonstrate the absence of non-spliced fragments; clarifying the 1/12 probability argument (or moving it to the Discussion with transcript-number context); providing sequences for the fluorophore variants and fusion tags; and minor prose corrections ("Regardless if" → "Regardless of whether"; "dependent on three things" → "dependent on three factors"). Reviewer 2 noted that the gene-size limitation, though tested, is not explicitly discussed in the main text.

    6. Author response:

      We are pleased that the reviewers viewed the core demonstration (that Dscam mutually exclusive splicing is preserved in a vector and that exon alternates can be replaced with genes of interest) as a solid foundation for the system. We agree that the manuscript would be strengthened by clearer quantitative characterization of expression, additional controls for fluorophore imaging, improved presentation of the figures, and more precise wording about the current scope of evidence. In a revised manuscript, we plan to address these points by adding or clarifying quantitative expression analyses, including S2 cell validation data, adding appropriate imaging controls where available, revising claims about 12-transgene expression to distinguish design capacity from direct experimental demonstration, and improving figure labels and legends throughout.

      We also plan to expand the discussion of PXGS limitations, including cell-type dependence on Dscam splicing machinery, possible position effects, and gene size considerations. Finally, we will improve Methods reporting by adding resource identifiers, cell culture quality control information, statistical design details, and data/code availability statements where appropriate.

      We appreciate the opportunity to revise the manuscript and believe these changes will make the strengths and limitations of PXGS clearer to readers.

    1. eLife Assessment

      This study presents a valuable finding on linking the frequency of neural activity to cortical depths of blood flow in a naturalistic setting of participants listening to music. The presentation of evidence in the version of the original submission is incomplete, as further clarifications in methods and results, as well as performing additional analyses, would strengthen the study. The work will be of interest to cognitive neuroscientists working on multimodal recordings, auditory perception and music.

    2. Reviewer #1 (Public review):

      Summary:

      In their submitted work, Lee and colleagues examine the correlation between electrophysiological activity as measured by SEEG, and layer-specific activation patterns, as measured through 7T fMRI. This analysis was performed using patients undergoing monitoring for epilepsy surgery guidance, as well as healthy controls, as they both listened to music.

      They find that, in general, higher-frequency SEEG activity correlated positively with the fMRI signal, while lower frequencies correlated negatively. Across cortical depth, higher-frequency activity correlated positively with middle-to-upper layers, whereas lower frequencies showed their strongest negative correlations in superficial layers.

      Strengths:

      This is an interesting physiological study in that, to the best of this reviewer's knowledge, it has not been done before with auditory stimuli using the combination of iEEG (as opposed to scalp EEG) and fMRI. The framework fits well with models of layer-specific feedforward versus feedback processing (e.g., Bastos et al., 2012).

      Weaknesses:

      Its main limitations are a lack of specificity to the acoustic stimuli, the absence of correction for venous draining, and the fact that it is largely a replication/port of prior work.

    3. Reviewer #2 (Public review):

      Summary:

      The authors present an investigation of the relationship between the iEEG frequency bands signal and hemodynamic responses at different cortical depths. Based on this, the authors aim to uncover the layered origin of iEEG signals at different frequencies. The authors then interpret their results in terms of feedforward and feedback processing, arguing that the correlations between fMRI and iEEG signals reflect the interaction between both processes. In addition, the authors aim to infer the extent to which these processes are involved during naturalistic music processing.

      Strengths:

      This study combines the neural recording methodologies yielding the highest spatio-temporal precision achievable in humans, while using naturalistic auditory stimuli. This combination of recording methods and experimental design offers key insights regarding the precise origin of iEEG signals, which is necessary to improve the interpretability of future iEEG studies.

      Weaknesses:

      (1) The current framing of the paper leads the authors to interpret their findings in ways that are not warranted by the data. The main analysis of the paper consists of correlating the hemodynamic responses from different layers with iEEG signals from different frequency bands, which enables us to infer the relationship between the two signals. It does not, however, enable us to draw inferences regarding the extent of feedforward and feedback processing and the interaction between the two during naturalistic auditory processing. This would require comparing hemodynamic responses in different cortical layers or frequency bands activation against some baseline condition. Based on the presented analysis, statements such as "our frequency-specific results demonstrate that naturalistic music perception seamlessly integrates both feedforward and feedback processing streams" should be removed.

      (2) The presentation of existing literature omits key details and findings, making it difficult to fully understand the research question the authors are trying to address. For example, the author mentions studies showing that feed-forward and feedback processing are segregated across cortical layers and that feed-forward and feedback processing have distinct time-frequency signatures (lines 47-58). However, the authors do not mention which cortical layer or which frequency band is associated with which kind of processing. As a result, it is difficult for the reader to determine what exact hypothesis the author is trying to test in the study. This might also relate to the confusion raised in (1).

      (3) The method section omits key details. When describing the paradigm, the authors do not describe how the tones were presented, nor how the signals were synchronized. Similarly, there is no mention of the pipeline used for iEEG electrodes localization. In addition, the exact regressors that entered the generalized linear model of hemodynamic responses are not clearly stated: were all regressors (frequency bands + acoustic signal + HFA) entered together in a single model or in separate models? The mention of a cubic spline is also not sufficient for the reader to understand what was done and for which purpose. Finally, the exact tests used for some comparisons are omitted (in Figure 2, for example, no mention of the exact test used to compare betas between A1 and A2). The current structure of the method section is also quite difficult to follow: the authors switch back and forth between describing acquisition protocols and participant counts, for example.

      (4) The lack of methodological details (as described in point 3 above) casts doubts about the validity of some of the statistical tests reported. Throughout the paper, the authors present quantitative statements and statistical tests comparing the fitted beta parameters between brain regions (A1 and A2) and cortical depths. However, the authors do not mention any normalization procedure taken to ensure that the scale of the signals being compared was equated. If the overall magnitude of the signals in A1 differs from that of A2, the mean of the beta distribution is expected to differ as well. Similarly, if the signal-to-noise ratio differs between brain regions or cortical layers, so should the variance of the beta parameters across subjects, which might break the homoscedasticity assumption of some tests, which might or might not be a problem depending on the exact test the authors used (hence the importance of reporting them).

    1. eLife Assessment

      This study provides a useful anatomical resource by mapping the expression of four putative chemoreceptors in spinal cerebrospinal fluid-contacting neurons (CSF-cNs) of larval zebrafish. These descriptive findings offer an interesting entry point to explore how the nervous system senses signals within the spinal fluid microenvironment. The evidence supporting the spatial expression patterns of these receptors is convincing, utilizing high-resolution hybridization chain reaction (HCR) to validate previous transcriptomic data. However, the evidence remains incomplete regarding the actual functional roles of these receptors, as the study lacks protein-level validation, evidence of ligand availability in the CSF, or functional assays to demonstrate active chemoreception.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript examines the expression of putative chemoreceptors in CSF-contacting neurons of the larval zebrafish spinal cord. Using in situ hybridization, the authors show that sstr2a is preferentially expressed in ventral CSF-cNs, whereas grm2a, ptprna, and ldlrad2 are detected in both ventral and dorsolateral CSF-cNs, with additional expression in neighboring cells around the central canal.

      Strengths:

      The study provides useful anatomical information on the expression of putative chemoreceptors in CSF-contacting neurons. The experiments appear to be carefully performed, and the results are clearly presented with high-quality illustrations and informative schematics.

      Weaknesses:

      This work remains largely descriptive and based on mRNA expression. Therefore, the proposed roles in chemoreception, ligand sensing, lipid capture, or long-range CSF signaling remain speculative without protein-level or functional validation.

    3. Reviewer #2 (Public review):

      Summary:

      Verran et al. leverage a previously published RNAseq dataset of zebrafish cerebrospinal fluid contacting neurons (CSF-cNs) to identify potential receptors involved in chemosensory signalling in these neurons. They then validate expression of the identified receptors by hybridization chain reaction (HCR) in zebrafish larvae. This way they uncover potential roles for the somatostatin receptor Sstr2a, metabotropic glutamtate receptor Grm2a, LDL receptor Ldlrad2 and the Phosphatase receptor Ptprna, suggesting the existence of numerous chemo-sensory pathways in CSF-cNs and providing a potential entry point for further investigation.

      Strengths:

      This is a useful resource; the provided HCR data that demonstrates expression of these receptors in CSF-cNs is convincing, and the finding that CSF-cNs express these receptors is interesting.

      Weaknesses:

      The overall insight provided by this manuscript is rather limited, essentially just demonstrating the expression of 4 receptors in CSF-cNs, whose expression was predicted to be enriched in these neurons anyway by a previously published dataset.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to identify new molecular pathways that could enable long-range signaling through the cerebrospinal fluid (CSF), focusing on a specialized class of neurons called CSF-contacting neurons (CSF-cNs) in larval zebrafish.

      Strengths:

      Anatomical validation of transcriptomic candidates using HCR, providing high-resolution spatial mapping of chemoreceptor expression in CSF-contacting neurons and neighboring spinal cord cells. The work broadens the potential understanding of CSF-cNs and offers a resource for future functional investigations of CSF-mediated signaling.

      Weaknesses:

      The principal limitation of the study is that the conclusions remain largely transcriptomic and inferential. Although HCR convincingly validates mRNA expression, no protein-level evidence is provided to demonstrate receptor translation or subcellular localization, leaving uncertainty regarding functional receptor availability at the CSF interface. Moreover, the study does not establish whether the proposed ligands are present in the relevant CSF microenvironment or engage the identified receptors in vivo. As such, the functional significance of the proposed chemosensory pathways remains speculative.

    1. eLife Assessment

      This work provides a reassessment of VBIT-4, a compound previously proposed to inhibit oligomerization of the crucial protein known as the mitochondrial voltage-dependent anion channel. Combining complementary experimental approaches with molecular dynamics simulations, the authors provide compelling evidence that VBIT-4 primarily disrupts lipid membranes and induces channel-independent cytotoxicity. The study has important implications for interpreting previous work using VBIT-4 as a probe of channel function and highlights the need to consider membrane-disruptive effects when evaluating drug mechanisms.

    2. Reviewer #1 (Public review):

      Summary:

      The Voltage-Dependent Anion Channel 1 (VDAC1) is the most abundant β-barrel protein in the outer mitochondrial membrane and the main conduit for metabolite and ion exchange between the cytosol and mitochondria. Its oligomerization has been proposed to control mitochondrion-mediated apoptosis, making it a prime target for therapeutic intervention in diseases associated with excessive cell death, such as neurodegenerative disorders and autoimmunity. VBIT-4 is a small molecule developed to inhibit VDAC oligomerization and has shown therapeutic potential in various preclinical models. Despite its widespread use, the mechanism of action of VBIT-4 has not yet been fully elucidated. In this paper, Ravishankar et al. combine a suite of biophysical approaches with computer simulations to demonstrate that VBIT-4 forms water-permeable defects in membrane bilayers without any detectable effects on VDAC1 channel properties or oligomerization. Furthermore, cytotoxicity assays revealed identical VBIT-4 IC50 values in wild-type and VDAC1-KO cells, indicating that its activity does not depend on VDAC1. Collectively, these findings cast significant doubt on the widely held assumption that VBIT-4 is a specific inhibitor of VDAC1 oligomerization. Instead, it appears that VBIT-4 functions as a membrane-active compound.

      Strengths:

      This is a carefully conducted and well-written study that highlights potential side effects of VBIT-4, a compound that has been used to study the role of VDAC1 in a range of physiological and pathological conditions. The work is of interest to a broad readership by showcasing the importance of a systematic assessment of drug-membrane interactions to identify potential off-target membrane-driven effects of small molecules that may be mistakenly attributed to the inhibition of specific proteins. Its strength lies in the variety of complementary approaches the authors used to rigorously challenge the effect of VBIT-4 on VDAC1 organization and function. Overall, the experimental data are compelling and of high quality.

      Weaknesses:

      The authors used high-speed atomic force microscopy (HS-AFM) to study the impact of VBIT-4 on VDAC1 oligomerization in real time at nanoscale resolution. Toward this end, they adsorbed POPC:POPE:cholesterol membranes reconstituted with or without VDAC1 on mica. This revealed that the addition of VBIT-4 produced small perforations in the bilayer that were independent of VDAC1. In the absence of VBIT-4, VDAC1 showed the characteristic honeycomb topography that the authors described in a previous study (Reference 17). To quantitatively assess whether VBIT-4 affects VDAC1 organization, they analyzed protein compaction within clusters using inter-protein distance measurements. This analysis revealed no significant difference in VDAC1 organization between control conditions, 1 uM and 10 uM VBIT-4, supporting a model in which VBIT-4 primarily perturbs the lipid matrix rather than VDAC1 assemblies. This conclusion is based on the assumption that VDAC channels retain some lateral mobility in bilayers adsorbed onto mica. Do the authors have evidence that this is indeed the case? Did they also perform HS-AFM on VDAC1-containing membranes treated with VBIT-4 prior to adsorption onto mica?

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript challenges the widely used interpretation of VBIT-4 as a specific inhibitor of VDAC1 oligomerization, arguing instead that it acts primarily as a membrane-active compound. Using high-speed atomic force microscopy, electrophysiology, liposome leakage assays, Laurdan fluorescence, microscale thermophoresis, coarse-grained molecular dynamics, and cell-based assays in wild-type and VDAC1-knockout HeLa cells, the authors show that VBIT-4 partitions into lipid bilayers, induces membrane defects and leakage, and causes VDAC1-independent cytotoxicity within a concentration range commonly used to infer VDAC1-specific effects.

      Strengths:

      The main strength is the convergence of several independent approaches to the same conclusion. Atomic force microscopy directly visualizes VBIT-4-induced defects in lipid regions while VDAC1 assemblies remain apparently intact. Electrophysiology separates VDAC1 channel behavior from background membrane conductance and shows that VBIT-4 does not measurably alter VDAC1 conductance or voltage gating, while increasing nonspecific membrane permeability. Lipid-only membranes, lipid nanodiscs lacking VDAC1, and VDAC1-knockout cells provide important controls supporting a VDAC1-independent mechanism.

      The wild-type versus VDAC1-knockout cytotoxicity comparison is a particularly strong test of VDAC1 independence. The observations that VBIT-4 is poorly soluble, aggregation-prone, and sensitive to storage conditions are also important, as they offer a plausible explanation for variability across previous studies. The revised manuscript is further strengthened by quantitative analysis of VDAC1 organization in atomic force microscopy images and by simulations including multiple VDAC1 molecules.

      Weaknesses

      The main limitation is that the conclusion that VBIT-4 does not affect VDAC1 oligomerization is strongest for the specific readouts used here: atomic force microscopy measurements of cluster compaction, VDAC1 channel properties, and simulated assembly behavior. These are direct and informative measurements, but they are not identical to the chemical cross-linking readouts used in much of the prior VBIT-4 literature. Readers should therefore distinguish between VDAC1 cluster organization in membranes, as measured here, and cross-linking-defined VDAC1 proximity.

      A second limitation is the uncertainty around effective VBIT-4 concentration. Because VBIT-4 is poorly soluble, aggregation-prone, pH-dependent, membrane-partitioning, and storage-sensitive, nominal added concentration may differ substantially from the concentration of active compound available in each assay. This complicates comparisons across the different in vitro, simulation, cellular, and previously published assays.

      The coarse-grained simulations provide a coherent mechanistic framework for membrane partitioning, aggregation, and defect formation. However, the VBIT-4 coarse-grained model is newly parameterized and is used to support a quantitative partitioning argument. The manuscript would be easier to interpret if the coarse-grained-derived partition coefficient were reported with uncertainty, convergence information, and protonation state, and compared with a matched all-atom octanol-water partition estimate from the same atomistic model used to build the coarse-grained mapping. This matters because the partitioning argument is used quantitatively to relate micromolar aqueous VBIT-4 to millimolar concentrations in the bilayer.

      Finally, the cellular data strongly support VDAC1-independent cytotoxicity, but the lower-dose mitochondrial functional phenotypes were not directly compared between wild-type and VDAC1-knockout backgrounds. VDAC1 independence is therefore more directly established for cytotoxicity than for the lower-dose mitochondrial phenotypes.

      Overall, this work provides a valuable and timely reassessment of VBIT-4, and its central conclusion will be useful for researchers interpreting studies that use this compound as a probe of VDAC1 function.

    1. eLife Assessment

      This valuable study investigates the mechanisms underlying inter-item biases in visual working memory. By experimentally manipulating the relative noise levels of target and non-target items, the authors report bias patterns that are broadly consistent with predictions of their previously proposed normative demixing theory. However, the supporting evidence remains incomplete, as the manuscript lacks a sufficient description of the underlying theory, key assumptions, and a quantitative link between the model and behavioral data. The manuscript would be substantially strengthened by clearer exposition and stronger tests, including analyses of the full error distributions and comparisons with alternative models, which would increase its potential interest to the cognitive neuroscience and computational cognitive science communities.

    2. Reviewer #1 (Public review):

      Summary:

      Many previous studies have reported inter-item biases in visual working memory tasks. These biases can be either attractive or repulsive, depending on the particular experiments. It has been difficult to explain these biases in a unifying theoretical framework. Recently, Chetverikov (the first author of the current manuscript) proposed a demixing model for explaining these biases in Ref 22. That paper shows that both attractive and repulsive biases could emerge in the demixing framework depending on the noise properties. The current manuscript seeks to test the predictions of the demixing model experimentally in a series of new experiments and find evidence supporting the demixing model.

      Because previous modeling results described in reference 22 (which is a preprint) are essential in interpreting the results reported in the current manuscript, I also studied that preprint and used the results reported in that paper to help interpret the results in this paper. My comments below will also contain discussions of that modeling paper.

      Strengths:

      Overall, the computational model tested in the paper is novel and interesting.

      The demixing framework represents an appealing hypothesis that deserves further investigation.

      The current paper provides new empirical data showing that the target stimuli with the same absolute noise level can be either repelled from or attracted to non-target items, depending on the relative noise levels. The observation that biases depend on the relative noise levels is by itself an interesting one, and is consistent with the prediction of the demixing model.

      Weaknesses:

      While this manuscript contains interesting new experimental observations and theoretical ideas, it has several substantial problems in its current form, which limit the conclusions that can be drawn. The description of the computational model is too brief. The key modeling assumptions need to be better motivated and explained. As the computational models generate different predictions in different regimes, it is a bit difficult to evaluate how well the experimental data support the model at a more quantitative level. Also, the results focused on studying the biases in the behavior; it is unclear whether the model can fully explain the behavior data (such as error distributions or behavioral precision).

      Major concerns:

      (1) Concerns/suggestions regarding the computational modeling

      The current paper seeks to test the predictions of the demixing-based computational model proposed in reference 22. There are several problems with the modeling component in the current paper.

      (1a) The description of the model is too brief and difficult to understand. Although the model was proposed in reference 22, it would still be beneficial to provide more details of the model so that readers can understand and appreciate the strengths/limitations of the model.

      The generative model and the inference procedure could be better explained to better link the model to the behavior. In particular, how was the observer's behavioral report in each trial modeled? This requires more explanation because currently the demixing procedure estimates four parameters for a given trial, yet for a given trial, only one behavioral report was produced (e.g., current Experiment 1), or two reports were produced sequentially (e.g., current Experiment 2).

      (1b) Key modeling assumptions need better justification.

      One such key assumption is that on a given trial, each stimulus triggers many samples (or approximately, an entire response distribution), rather than a single sample. This assumption deviates substantially from prior work on ideal observer models. It was not clear whether this assumption is realistic. For the type of stimuli used in the current experiments, perhaps one can argue that each pixel corresponds to one sample of brain activity, thus collectively each stimulus should trigger many samples of activity in the brain. If this were to be the case, it would have two implications. First, the noise parameter in the model should be directly related to the magnitude of the stimulus noise. Thus, one should be able to plug these experimentally-controlled parameter values into the model to directly generate predictions about the biases. Second, when using stimuli with no stimulus variability (e.g., simple grating stimuli), the predicted biases should change. However, it wasn't clear whether this would hold experimentally, i.e., using gratings would lead to different biases or no biases.

      If the variability of the samples for a given stimulus involves neural noise, it would be useful to justify why it is reasonable to consider that many samples were generated per stimulus.

      (1c) As mentioned in (1b), the model assumes that on each trial, a large number of samples was generated. It would be useful to study and report how the prediction would change when the number of samples generated per stimulus is small. In particular, what happens when each stimulus only generates one measurement? This might be useful for interpreting previous experiment results with grating stimuli.

      (1d) Reference 22 studies how the predicted biases depend on the d-prime of the identifying dimension and found that the pattern of the biases varies substantially depending on the information available for the identifying dimension. However, the current paper didn't really discuss this important point. It is also unclear what parameters the authors used for the d-prime of the identifying dimension. Was it fitted directly to the data? The Methods section has some description on the "identifiability dimension", but it was a bit obscure.

      Intuitively, when the d-prime of the identifying dimension is very large, the demixing problem becomes irrelevant. In this case, there should not be any biases induced by demixing. In the case of the d-prime for the identifying dimension is 0, the problem should reduce to the simplified 1-d problem studied in reference 22. If my reading of reference 22 was correct, they reported different conclusions. It would be useful to clarify these points.

      In any case, the d-prime of the identifying dimension appears to be a key parameter. It would be great to constrain this parameter using the empirical data. When the d-prime of the identifying parameter is small, the observer would easily confuse the probed stimulus with the other stimulus in a given trial. This should lead to poor task performance. Thus, it may be possible to directly estimate the value of the d-prime of the identifying dimension based on the observer's performance, and then use this parameter to generate model predictions accordingly.

      (1e) The current model assumes that a large number of samples are generated per stimulus and the brain can manipulate this information to perform the demixing task. It was well documented that visual working memory has a capacity limit (i.e., it can only hold information about a few items); this discrepancy needs to be clarified or addressed.

      (2) How well the computational model can explain the experimental data remains not entirely clear

      The authors show that there exists a parameter regime that can qualitatively explain the experimental finding. They also show that it is possible to fit the model to the data to explain the bias patterns. However, given that the model is flexible, it would be stronger if the authors could show that the same parameters that explain the biases could also explain other aspects of the behavior, for example, the magnitude of the errors.

      In other words, the model is not well constrained in the way it was tested in the paper. But it should be possible to improve it. First, if the noise parameter in the model is determined by the stimulus variability, one can determine it directly based on the external noise in the stimuli (discussed also in 1b) and see what prediction it leads to. Second, from the behavioral data, it may be possible to estimate the noise for the identifying dimension. Doing so will help better constrain the model.

      It would also help if the authors could report the best-fitted parameters from the experimental data. From these parameters, one can simulate synthetic data and apply the demixing model to see if the error distribution of the simulated observers is indeed similar to the experimentally measured error distribution. That way, one can check whether the fitted parameter explains the observer's behavioral performance beyond the biases.

      Other comments:

      (1) How does the model account for the swap errors? I am not sure I understood the way how the swap errors were treated in the paper. To me, substantial swap errors seem to be a consequence of having low d-prime values for the identifying dimension; that is, if there is only little information to discriminate the identity of the two stimuli, swap errors would be large. However, this possibility didn't seem to be mentioned in the paper.

      (2) Since the solution of the demixing problem was obtained using a numerical procedure based on EM. It would be useful to check whether the initialization has affected the biases obtained.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the origins of inter-item biases in visual working memory. The authors proposed a computational model where overlapping memory signals are disentangled, inducing memory biases that depend on relative noise levels across items. The key theoretical advance is the prediction that bias direction depends not only on absolute memory noise but on the relative noise levels of target and non-target representations. Using four experiments with color mosaics whose color variability manipulates memory precision, the authors report that biases reverse as a function of relative noise in a manner predicted by the model.

      Strengths:

      The manuscript is clearly written and theoretically motivated. The experiments are well designed and provide converging evidence for a distinctive and non-intuitive prediction of the proposed model. I found the central result compelling: independently manipulating target and non-target noise leads to qualitatively different bias patterns, consistent with the model's prediction that relative noise is a key determinant of bias direction.

      Weaknesses:

      The main limitation is that the evidence establishes consistency of the data with the proposed Demixing Model, but does not demonstrate that the model provides a unique explanation of the data. Although the manuscript argues that dominant theories struggle to account for the observed reversals, no formal comparison with alternative computational frameworks is presented. In addition, model fitting results are reported only briefly, making it difficult to evaluate fit quality at the level of individual observers.

    4. Author response:

      Reviewer #1 (Public review):

      Strengths:

      Overall, the computational model tested in the paper is novel and interesting.

      The demixing framework represents an appealing hypothesis that deserves further investigation.

      The current paper provides new empirical data showing that the target stimuli with the same absolute noise level can be either repelled from or attracted to non-target items, depending on the relative noise levels. The observation that biases depend on the relative noise levels is by itself an interesting one, and is consistent with the prediction of the demixing model.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      While this manuscript contains interesting new experimental observations and theoretical ideas, it has several substantial problems in its current form, which limit the conclusions that can be drawn. The description of the computational model is too brief. The key modeling assumptions need to be better motivated and explained. As the computational models generate different predictions in different regimes, it is a bit difficult to evaluate how well the experimental data support the model at a more quantitative level. Also, the results focused on studying the biases in the behavior; it is unclear whether the model can fully explain the behavior data (such as error distributions or behavioral precision).

      We agree that the model description should be expanded and that quantitative agreement with the data should be assessed more thoroughly, and we plan to address this during the revision. In the initial version of the manuscript, we aimed to highlight the qualitative agreement of the data with the novel and counterintuitive predictions by the model. While the reviewer is correct that the model "generates different predictions in different regimes," the particular predictions we test (the interaction between the target noise level and the parity of the target and non-target noise levels in Experiments 1-3, and the effects of non-target noise when the target noise is held constant in Experiment 4) hold across regimes (Figure S1 shows this for the former prediction). We aim to further expand on this point in the revision.

      Major concerns:

      (1) Concerns/suggestions regarding the computational modeling

      The current paper seeks to test the predictions of the demixing-based computational model proposed in reference 22. There are several problems with the modeling component in the current paper.

      (1a) The description of the model is too brief and difficult to understand. Although the model was proposed in reference 22, it would still be beneficial to provide more details of the model so that readers can understand and appreciate the strengths/limitations of the model.

      The generative model and the inference procedure could be better explained to better link the model to the behavior. In particular, how was the observer's behavioral report in each trial modeled? This requires more explanation because currently the demixing procedure estimates four parameters for a given trial, yet for a given trial, only one behavioral report was produced (e.g., current Experiment 1), or two reports were produced sequentially (e.g., current Experiment 2).

      We will provide more details about the model and how it was fitted to the data. Please note that the model parameters were fitted per subject and condition, not per trial: 2 hue noise parameters,  and , corresponding to the noise of the target and non-target item across 4 noise combination conditions, plus a shared identifiability noise, , across conditions, determining the discriminability of the items along the identifying dimension. This strongly limits model flexibility as only 3 parameters (including the shared  across conditions) are used to create the bias curve for each subject in each condition.

      (1b) Key modeling assumptions need better justification.

      One such key assumption is that on a given trial, each stimulus triggers many samples (or approximately, an entire response distribution), rather than a single sample. This assumption deviates substantially from prior work on ideal observer models. It was not clear whether this assumption is realistic. For the type of stimuli used in the current experiments, perhaps one can argue that each pixel corresponds to one sample of brain activity, thus collectively each stimulus should trigger many samples of activity in the brain. If this were to be the case, it would have two implications. First, the noise parameter in the model should be directly related to the magnitude of the stimulus noise. Thus, one should be able to plug these experimentally-controlled parameter values into the model to directly generate predictions about the biases. Second, when using stimuli with no stimulus variability (e.g., simple grating stimuli), the predicted biases should change. However, it wasn't clear whether this would hold experimentally, i.e., using gratings would lead to different biases or no biases.

      If the variability of the samples for a given stimulus involves neural noise, it would be useful to justify why it is reasonable to consider that many samples were generated per stimulus.

      We are grateful to the reviewer for raising this point, and we will provide more details on it in the revision. In brief, we believe that it is the standard ideal observer assumption of one sample per trial that is unrealistic and works only in cases when there is a single signal source, so that the samples can be simplified to a single average. Consider that determining a stimulus value is a similar problem for an ideal observer to the one that a researcher who aims to decode neural data from populations of neurons (or fMRI voxels) has to solve. Different populations of neurons would provide responses that match different stimuli – in essence, creating different samples in an ideal observer framework. Thus, even without external noise, the demixing problem would be present when there is more than one stimulus, but internal noise is much more difficult to control, so in our experiments, we used multi-colored stimuli.

      (1c) As mentioned in (1b), the model assumes that on each trial, a large number of samples was generated. It would be useful to study and report how the prediction would change when the number of samples generated per stimulus is small. In particular, what happens when each stimulus only generates one measurement? This might be useful for interpreting previous experiment results with grating stimuli.

      This is an interesting point that we aim to address in the revision.

      (1d) Reference 22 studies how the predicted biases depend on the d-prime of the identifying dimension and found that the pattern of the biases varies substantially depending on the information available for the identifying dimension. However, the current paper didn't really discuss this important point. It is also unclear what parameters the authors used for the d-prime of the identifying dimension. Was it fitted directly to the data? The Methods section has some description on the "identifiability dimension", but it was a bit obscure.

      Intuitively, when the d-prime of the identifying dimension is very large, the demixing problem becomes irrelevant. In this case, there should not be any biases induced by demixing. In the case of the d-prime for the identifying dimension is 0, the problem should reduce to the simplified 1-d problem studied in reference 22. If my reading of reference 22 was correct, they reported different conclusions. It would be useful to clarify these points.

      We are grateful for the suggestion to expand the discussion of this point and will do so in the revision. The reviewer is correct that for very large d-prime in the identifying dimension, the demixing problem solution is trivial. However, the 2D case does not resolve to the 1D case when d-prime reaches zero. This is because the identifying dimension is still used to identify which item to report—unlike in the 1D case, when the reported dimension is the same as the identifying one. Consider what happens if the observer in our task does not remember at all which stimulus was left and which was right. It would report the other item in 50% of cases, leading to a strong attractive bias.

      In any case, the d-prime of the identifying dimension appears to be a key parameter. It would be great to constrain this parameter using the empirical data. When the d-prime of the identifying parameter is small, the observer would easily confuse the probed stimulus with the other stimulus in a given trial. This should lead to poor task performance. Thus, it may be possible to directly estimate the value of the d-prime of the identifying dimension based on the observer's performance, and then use this parameter to generate model predictions accordingly.

      We apologize for the confusion. We constrain the discriminability of items in the "identifying" dimension using the  parameter that determines the noise in that dimension for both items. The means in this dimension are fixed at an arbitrary value, as means and noise are interchangeable when considering discriminability. We will revise the description of the fitting procedure accordingly. Regarding the use of the same values in predictions, while possible, we prefer to keep predictions separate from fitting to avoid them becoming postdictions. The curves for the fitted model in Figure 2 already illustrate what the model predicts under the fitted parameter values.

      (1e) The current model assumes that a large number of samples are generated per stimulus and the brain can manipulate this information to perform the demixing task. It was well documented that visual working memory has a capacity limit (i.e., it can only hold information about a few items); this discrepancy needs to be clarified or addressed.

      We are grateful to the reviewer for raising this point, which we will address in the revised discussion. Briefly, we believe that the number of samples in the ideal observer model does not correspond directly to the working memory “slots”.

      (2) How well the computational model can explain the experimental data remains not entirely clear

      The authors show that there exists a parameter regime that can qualitatively explain the experimental finding. They also show that it is possible to fit the model to the data to explain the bias patterns. However, given that the model is flexible, it would be stronger if the authors could show that the same parameters that explain the biases could also explain other aspects of the behavior, for example, the magnitude of the errors.

      It would also help if the authors could report the best-fitted parameters from the experimental data. From these parameters, one can simulate synthetic data and apply the demixing model to see if the error distribution of the simulated observers is indeed similar to the experimentally measured error distribution. That way, one can check whether the fitted parameter explains the observer's behavioral performance beyond the biases.

      We are grateful to the reviewer for raising this point. We both agree and disagree with the reviewer here. The predictions reported come from an earlier paper describing the model (ref. 22). In our opinion, this represents a pure hypothesis-driven approach, where a prediction is formulated first and then tested with subsequently collected data. The model we test is normative, not descriptive; its goal is not to fit the data as closely as possible, but rather to make predictions about internal brain mechanisms. We do not suggest, for example, that demixing is the sole source of biases, so the resulting bias pattern might differ significantly from the predictions. That the model fits the data is, therefore, an additional bonus. At the same time, we agree that it is interesting to test whether the model can explain other parameters of the data. Note that our current fitting procedure was not geared toward this; we optimized the model to explain only the bias curve. In the revision, we aim to test whether the model can also explain the error variability.

      In other words, the model is not well constrained in the way it was tested in the paper. But it should be possible to improve it. First, if the noise parameter in the model is determined by the stimulus variability, one can determine it directly based on the external noise in the stimuli (discussed also in 1b) and see what prediction it leads to. Second, from the behavioral data, it may be possible to estimate the noise for the identifying dimension. Doing so will help better constrain the model.

      External noise accounts for only a portion of the total noise, as evidenced by behavioral errors. Even for a single item, the total noise consists of the amount of information the observer samples from the stimulus, the variability of these samples (external noise), and early (applied to each sample) and late (applied after integration) internal noise. Therefore, external noise alone might not constrain the model in the right regime. Regarding the identifying-dimension noise, as noted above, we do constrain it with the data. However, we aim to explore these points further in the revision.

      Other comments:

      (1) How does the model account for the swap errors? I am not sure I understood the way how the swap errors were treated in the paper. To me, substantial swap errors seem to be a consequence of having low d-prime values for the identifying dimension; that is, if there is only little information to discriminate the identity of the two stimuli, swap errors would be large. However, this possibility didn't seem to be mentioned in the paper.

      We apologize for the confusion. We will further clarify and perhaps reassess the treatment of swap errors in the revision. The model itself produces swap errors when the stimuli sources are misidentified.

      (2) Since the solution of the demixing problem was obtained using a numerical procedure based on EM. It would be useful to check whether the initialization has affected the biases obtained.

      Indeed, this is a valid point, and it's why we use a multi-initialization strategy. For each simulation of a single trial sample set (e.g., 100 random samples), we use a large number of initial points (50 in the initial submitted manuscript) to ensure the obtained EM solution is truly optimal. Additionally, we conduct a large number of simulated trials (10,000 for each parameter combination) to ensure the accuracy of the bias distribution we obtain.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the origins of inter-item biases in visual working memory. The authors proposed a computational model where overlapping memory signals are disentangled, inducing memory biases that depend on relative noise levels across items. The key theoretical advance is the prediction that bias direction depends not only on absolute memory noise but on the relative noise levels of target and non-target representations. Using four experiments with color mosaics whose color variability manipulates memory precision, the authors report that biases reverse as a function of relative noise in a manner predicted by the model.

      Strengths:

      The manuscript is clearly written and theoretically motivated. The experiments are well designed and provide converging evidence for a distinctive and non-intuitive prediction of the proposed model. I found the central result compelling: independently manipulating target and non-target noise leads to qualitatively different bias patterns, consistent with the model's prediction that relative noise is a key determinant of bias direction.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      The main limitation is that the evidence establishes consistency of the data with the proposed Demixing Model, but does not demonstrate that the model provides a unique explanation of the data. Although the manuscript argues that dominant theories struggle to account for the observed reversals, no formal comparison with alternative computational frameworks is presented. In addition, model fitting results are reported only briefly, making it difficult to evaluate fit quality at the level of individual observers.

      We agree and we aim to provide a comparison with alternative models and an expanded description of the fitting results in the revision. Note, however, that the majority of existing models are descriptive, while we believe that as a normative model, the Demixing Model should be compared with other normative models, thus limiting the selection of competitors significantly.

    1. eLife Assessment

      This useful study reports on lifespan extension in C. elegans males that carry a mutation in a gene for an insulin receptor; while the observations are striking, the strength of evidence is currently incomplete. The central claim of male-specificity is undermined by the absence of direct hermaphrodite healthspan comparisons, and the reliance on a single mutant allele leaves open the possibility that background mutations, rather than daf-2 loss-of-function, drive the phenotype. Methodological details critical for reproducibility are also lacking, particularly regarding male housing density, censoring of plate-leaving animals, and the adequacy of replication for the key epistasis experiment. The work could be substantially strengthened by a targeted set of additional experiments and fuller engagement with the existing literature on sex-specific aging in C. elegans.

    2. Reviewer #1 (Public review):

      Summary:

      This paper provides interesting observations about the effects of a classical mutation in the daf-2 insulin-like receptor in male C. elegans. The observations are a contribution in and of themselves; however, the conclusions reached about these observations are not supported by the work presented. Most importantly, male-specific effects on healthspan measures are asserted without direct comparison to hermaphrodites. Perhaps more fundamentally, essential features of the methods and experimental design are lacking, which makes formal assessment of the results impossible, especially given our knowledge of negative male-male interactions, which have gone completely unacknowledged here. Indeed, there is a general lack of context for known sex differences in C. elegans, especially in terms of the core elements of longevity, which are presented here as entirely novel but in fact are not.

      Major comments:

      (1) The main overall criticism of the premise of the paper is that it lacks a clear hypothesis that would lead to explicit experimental tests. Instead, many of the results are observational, and the conclusions reached go beyond the actual experiments conducted. The goal should be explicit and consistent between the introduction/ discussion, and the findings should directly address the goal.

      The overall focus appears to be that daf-2 males have an extended lifespan for reasons that are different from hermaphrodites. This conclusion is apparently based on the observation of lipid reserves in mutant animals. However, none of the healthspan measures were conducted in parallel with identical measures in hermaphrodites. How can the authors then claim that males are unique? This is especially problematic since other studies have demonstrated that daf-2 hermaphrodites also have altered lipid composition (Vrablik 2015 Biochim Biophys Acta; Horikawa 2010 Mol Cell Endocrin).

      (2) The authors make unwarranted claims about causation from observational data that is correlative in nature. Again, they claim that male longevity is caused by increased lipid reserves. This may in fact be the case, but there is no evidence to show that this is causal, only that lipid reserves are increased in mutant animals. Causation requires an actual experiment, in this case, disrupting lipid maintenance in daf-2 males (e.g., Lapierre 2013 Autophagy). Their conclusions are consistent with their results, but their conclusions are much too strong given the nature of the evidence, especially given the concerns about proper comparisons to hermaphrodites.

      (3) With these concerns in mind, all conclusions related to male-specific effects should be statistically tested using a sex-by-treatment interaction term in the statistical model. This is obviously impossible for the healthspan data, but for lifespan, this can be directly tested using (genotype x sex interaction in the CPH analysis). Further, it is unclear why each of the replicates is shown separately in Figure 1.

      It is nice that the authors do not directly pool them, as most longevity studies do, but the replicate effects can be included in a more comprehensive model, which would yield an appropriate "average" effect curve.

      (4) There is an inadequate review of pre-existing literature and findings that predate the observations presented here. While this is not an issue in general, the authors present their work as entirely novel when it is not.

      In addition to Gems and Riddle (2000), which is tangentially cited in the discussion, the following papers should be cited and discussed in the introduction to clarify what is currently known and what remains to be explored:

      Partridge and Gems (2002) Mechanisms of aging: public or private?

      McCulloch and Gems (2007) Sex‐specific Effects of the DAF‐12 Steroid Receptor on Aging in Caenorhabditis elegans

      Hotzi et al (2018) Sex‐specific regulation of aging in Caenorhabditis elegans

      Al-Saadi et al (2025) Disruption of the insulin signaling pathway in C. elegans dramatically increases male longevity and enhances reproductive health late in life

      In addition, the authors assert that the study of sex differences is unstudied. If the authors are specifically referring to the sex differences in aging research, they should explicitly state that and revise their language to reflect that it is "understudied" rather than "unstudied". But as stated below, there are many studies that look at sex-specific differences in behavior, physiology, development, etc. This is most important in the context of sexual conflict, of which there are many studies that are directly relevant to the work presented here. The authors are encouraged to review some of these papers.

      (5) This is particularly important in the context of how the experiments presented here were actually conducted. The methods are inadequate to assess this, and the results would therefore be impossible to replicate in the absence of additional details. Exactly how many individuals were raised on each plate during the longevity assays (and other work) is critical to understanding the results of this study. This is because males have direct, chemically and physically mediated negative impacts on one another (see many papers from the Brunet and Murphy labs). Further, it is not even clear whether males and hermaphrodites were reared separately from one another. Males are known to leave plates without hermaphrodites, which requires appropriate inclusion of censoring criteria in studies such as these. It is unclear whether and how this was handled. Censoring is an essential feature of any longevity study and so needs to be explicitly described in the statistical methods.

      The methods describe the use of heat shock to induce the production of males, but it is unclear which generation is being used here. Ordinarily, males would be induced, and then male populations would be maintained by forced mating (picking to ensure that there is a high relative frequency of males) for several generations to eliminate any carryover effects of the heatshock itself. Were the heatshock males put directly into the longevity assays? If so, were hermaphrodites subject to identical treatment? This is confusing, and a potentially critical confound is not performed correctly.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript presents interesting observations regarding the exceptional longevity and improved healthspan of male daf-2 mutants. Given the comparatively limited focus on male aging in C. elegans, the study provides a potentially useful characterization of sex-specific effects associated with reduced IIS signaling.

      Strengths:

      The 4-fold increase in lifespan of male daf-2 mutants is a striking and unexpected observation. The altered fat metabolism between older daf-2 mutant males and hermaphrodites provides further evidence of sex-specific effects.

      Weaknesses:

      (1) A major limitation of the current study is that the conclusions rely primarily on a single daf-2 allele. It would strengthen the manuscript to validate at least the major observations using an independent daf-2 allele or through daf-2 RNAi. This is particularly relevant for the proposed male-specific enhancement of longevity and healthspan, as it remains unclear whether the observed effects broadly reflect reduced IIS signaling or may be influenced by allele-specific effects or background mutations.

      (2) The methods for male lifespan assays require additional detail. Although the authors state that males were generated and transferred every three days, it is not clear whether males were maintained singly or in groups, how many males were placed per plate, or how many were censored by fleeing. These details are particularly important for male aging assays, as male lifespan in C. elegans is known to be influenced by social interactions. Factors such as population density can affect survival and healthspan measurements. Clarifying these procedures would improve reproducibility and interpretation of the reported male-specific lifespan effects.

      (3) Because reduced IIS signaling in daf-2 mutants is known to alter metabolism, physiology, and potentially male-derived signaling, it would be interesting to determine whether the enhanced longevity of daf-2 males is influenced by altered male-male interactions or resistance to male-associated toxicity. In this context, clarification of whether lifespan assays were performed with grouped or individually maintained males would be valuable. If not already tested, lifespan analysis under isolated single-male conditions could help distinguish intrinsic longevity effects from potential contributions of population-dependent signaling or social interactions.

      (4) In Figure 2D, the body-length measurements in WT males appear somewhat unexpected, particularly the apparent increase between Day 14 and Day 20. Since adult worms are not typically expected to exhibit substantial growth at advanced ages, additional clarification regarding the measurement methodology would be helpful, including confirmation that the scale bars and image scaling were applied consistently across conditions.

      (5) The use of palmitic acid barriers following Beydoun et al. (2024) is appropriate; however, it would be helpful to clarify whether WT and daf-2 males exhibited comparable fleeing behavior under these assay conditions. Because male worms are highly prone to plate leaving and censoring, genotype-dependent differences in fleeing behavior could potentially influence survival analyses and the number of censored animals. In addition, as Beydoun et al. primarily characterized these barrier conditions using hermaphrodites, it would be useful to clarify whether comparable barrier effectiveness was observed in male lifespan assays.

      (6) Separately, Beydoun et al. (2024) also reported that palmitic acid barrier conditions can influence body-size measurements, whereas PEG-based barriers did not show similar effects on body size. It would therefore be useful to know whether comparable body-length trends were observed under alternative barrier conditions, particularly given the unexpected increase in WT male body length at later ages.

    4. Reviewer #3 (Public review):

      This manuscript reports a striking sex-specific effect of the daf-2(e1370) mutation on C. elegans lifespan. The authors show that male daf-2 mutants exhibit dramatically extended lifespan relative to wild-type males, wild-type hermaphrodites, and daf-2 hermaphrodites. The study also demonstrates increased lipid accumulation in these long-lived males, which is increased further over time, improved late-life motility, enhanced oxidative stress resistance, and a requirement for the downstream effector of daf-2, daf-16, for the longevity phenotype.

      The interest of the work is the magnitude and consistency of the lifespan effect. The authors report large increases in both median and mean lifespan in daf-2 males across independently replicated experiments. This is further supported by healthspan analyses and their finding that male daf-2 mutants maintain improved motility and stress resistance, which argues against the interpretation that lifespan extension merely reflects prolonged frailty. The genetic epistasis experiment demonstrating loss of the longevity phenotype in daf-2;daf-16 double mutants provides evidence that the effect depends on canonical insulin/IGF-1 signalling.

      The main limitation is that at least the first figure is rather an incremental increase on previous work examining the lifespan of daf-2 males, although the authors do indeed show that the effects can be much larger (or more 'plastic') than those previously published. While these findings are potentially important, the manuscript would certainly benefit from a more extensive discussion of how the results compare with prior studies of daf-2 mutants and male longevity, including possible explanations for the apparent discrepancies.

      The epistasis experiment shows that this exceptional longevity requires the expression of daf-16. However, in contrast to the initial experiments (Figure 1) that show three replicates of the lifespan experiment (the standard in lifespan work in this model), it appears that the daf-2;daf-16 experiment has only been performed once.

      In addition, the lifespan data for hermaphrodite daf-2 mutants appear somewhat unusual. Although the mean lifespan is increased, the median lifespan is reported to be only modestly greater than that of wild-type hermaphrodites. I know that this mutant can give lifespan curves that look like this, but either the use of another allele or of the experimental conditions and how these values compare with previously published daf-2(e1370) datasets would help readers interpret the magnitude of the male-specific effect.

      The lipid phenotype is intriguing. It would be interesting to expand this to examine somatic vs embryonic fat. In addition, I noted that in the methods section, the authors use palmitic acid to stop the male worms 'fleeing' the plates; is it possible to rule out the possibility that the daf-2 mutants are simply eating and metabolising/storing this fatty acid barrier differently than their wild-type counterparts? This would be worth considering and controlling for, particularly as male C. elegans have been shown to have dramatically altered metabolic transcriptional profiles. If indeed this increased lipid is responsible for the extreme longevity of the daf-2 mutant males, it would be desirable to try to link this mechanistically to the phenotype.

      Overall, the evidence convincingly supports the conclusion that male daf-2(e1370) mutants are exceptionally long-lived under the conditions tested and that this phenotype requires DAF-16. The work has the potential to make an important contribution to understanding sex-specific regulation of ageing, although further contextualisation within the existing literature would strengthen the manuscript.

    1. eLife Assessment

      This valuable study provides novel evidence that congenital aphantasia is associated with structural differences in frontotemporal and cingulate systems, with relative sparing of early visual regions and major visual pathways. The multimodal structural imaging approach is carefully implemented and will be of interest to researchers studying mental imagery and aphantasia. However, the strength of evidence is incomplete because the data cannot adjudicate between alternative cognitive interpretations, and the multiple discovery streams make the findings better viewed as key exploratory evidence, rather than as establishing a definitive structural phenotype of aphantasia.

    2. Reviewer #1 (Public review):

      Summary

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (3) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

    3. Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths:

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (4) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

      We will add two recent in-press Cortex papers to the Discussion. One provides lesion-based double-dissociation evidence against V1 as a necessary causal substrate of visual imagery. The other shows that aphantasic individuals can display visualizer-like oculomotor patterns during mental map exploration despite reporting little or no imagery vividness. Together, these studies help clarify our interpretation of our null V1 findings and structural effects in higher-order brain regions, which are consistent with aphantasia involving altered integration or access rather than a primary V1-dependent imagery deficit.

      Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

      We thank the reviewer for this thoughtful and constructive comment. Lack of introspective report of voluntary imagery is arguably the defining signature of aphantasia. This motivated us to primarily interpret our anatomical findings in a broader cognitive context of higher-order control, internal monitoring, and conscious access in aphantasia. We expect that a reliable behavioural test measuring imagery sensitivity and accessibility would allow us to direct link these findings to individual imagery ability. Nevertheless, to our best knowledge, this kind of test on imagery is still missing. Instead, our findings point to some plausible structural signature or brain regions that may be related to conscious imagery, which motivate future studies to examine their direct or causal roles. We agree with the reviewer, future studies should test the relationship between these anatomical structures and the accessibility to internal representation, together with related aspects of internal monitoring. We will therefore add a dedicated paragraph to discuss the plausible cognitive mechanisms during the revision.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

      We are grateful to the Reviewer for the constructive and thoughtful assessment of our manuscript. In response to the reviewer’s comments, we will revise the manuscript to clarify the dependency between diffusion-derived analysis streams, to state more explicitly the biological limits of diffusion MRI metrics, to add a null-network sensitivity analysis for the clustering coefficient findings, to include confidence intervals for reported effect sizes, and to temper the interpretation of the cortical thickness result. We will also revise the Abstract and Discussion to better reflect the exploratory nature of the study and to frame the findings as hypotheses requiring replication in larger independent samples. We believe that these revisions will make the manuscript more balanced, transparent, and appropriately cautious, while preserving the central conclusion that congenital aphantasia is associated with structural differences centered on higher-order frontotemporal and cingulate systems.

    1. eLife Assessment

      This valuable study identifies Gcn5 as a regulator of blood cell development in the Drosophila lymph gland, with links to autophagy and nutrient-sensing mTORC1 signalling. The evidence is solid that altering Gcn5, autophagy genes and mTORC1 activity perturbs blood cell homeostasis, and the revised manuscript adds helpful genetic and quantitative analyses. However, the evidence for a clean linear Gcn5-mTORC1-TFEB/autophagy pathway is insufficient, because several cell-type-specific phenotypes remain difficult to reconcile and the pathway logic relies on different genetic tools, cell populations and pharmacological perturbations.

    2. Reviewer #1 (Public review):

      In their manuscript Arjun et al. investigate the role of the histone acetyl transferase Gcn5 in controlling drosophila blood cell homeostasis in the larval lymph gland. Using gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland cell populations, they show that these manipulations impact (but in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation, DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and show that reducing the expression of the autophagy machinery affect blood cell differentiation. By using drugs as well as genetic approaches to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR, but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      Overall, the main conclusions are sound but interpreting several lines of experiments and results remain complicated. Consequently, the overall picture of the role of Gcn5 in Drosophila larval lymph gland development, and its relationship to mTOR and autophagy, remains unclear.

    3. Reviewer #2 (Public review):

      Summary:

      Drosophila haematopoiesis has been shown to be governed by a number of signalling pathways such as JAK/STAT and Dpp. This important study shows a role for nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it TFEB activity thus acting as a fine-tuning mechanism which maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy has a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weakness:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections do not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. And the modulations of Gcn5 in PSC also has variable effects across hemocyte subpopulations which is not explored in the manuscript. Interestingly, also the domain deletion constructs show differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained. While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4. The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      Comments on the revised manuscript:

      Overall, the revisions make the narrative more coherent. The authors have also added data which substantiates their conclusions.

      However, in some instances, the authors are not clearly able to explain the discrepancies in the data (Gen-5 depletions under the Hml-Gal4 in the whole larval lysates remove p62 completely) which is not ideal.

      A query regarding the discrepancies in the immunofluorescence data: The authors have removed the IF data which suggested that there could be differences in the shuttling of Gcn5 between the nucleus and cytoplasm. The authors suggest that immunofluorescence issues are at the root of these variable results, but the reviewer wonders whether there could be further unexplored mechanisms re: shuttling that is unexplored here and would have been potentially novel.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In their manuscript, Arjun et al. investigate the role of the histone acetyltransferase Gcn5 in the control of drosophila blood cell homeostasis in the larval lymph gland. They use gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland populations to show that these modulations impact (in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation or DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and they show that decreasing the expression of the autophagy machinery increases blood cell differentiation. Using drugs to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      While the authors did a lot of experiments and good quantifications of the blood cell phenotypes, many results do not make much sense or do not bring valuable information about Gcn5 mode of action. Several conclusions of the manuscripts are not backed by solid data (e.g. that Gcn5 action is mediated by TFEB and the autophagy machinery) and different aspects of the literature are not well taken into consideration. Some results (such as the validation of the knockdown and overexpression of Gcn5) seem flawed. There are some concerns about the results obtained with gcn5 zygotic mutants and an interpretation of the phenotypes observed upon manipulation of Gcn5 expression in different cell types is missing.

      We have now performed several experiments to address the comments raised by the reviewer and have also provided possible explanation of the phenotypes in cases where it was lacking.

      Important revisions are needed to improve the quality of the manuscript and confirm the authors' findings.

      Reviewer #2 (Public Review):

      Summary:

      Drosophila hematopoiesis has been shown to be governed by a number of signaling pathways such as JAK/STAT and Dpp. This important study shows the role of nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it, TFEB activity thus acting as a fine-tuning mechanism that maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy have a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub-sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weaknesses:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections does not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. The modulations of Gcn5 in PSC also have variable effects across hemocyte subpopulations which is not explored in the manuscript.

      We have now explained and discuss why Gcn5 modulation could be affecting the PSC size. Please check Discussion section Paragraph 1 line 10 onwards.

      Interestingly, also the domain deletion constructs show a differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained.

      Currently, with our observations all that we can comment about that data is that expression of domain deletion mutants causes aberrant hematopoiesis indicating a dominant negative phenotype since they are expressed in the wild type genetic background. Beyond this, we will be exploring mechanistically how these domains are functioning during hematopoiesis in future studies. We have already described the dominant negative effect in the text: Discussion Section Paragraph 3.

      While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4.

      We have now revised Figure 1 and have only included the images with Collier-Gal4 and Dome-Gal4 which clearly shows both the niche cells, Dome-positive progenitors and Dome-negative cells of the primary LG lobe essentially showing that Gcn5 is expressed throughout the primary LG lobe. In Fig. 1C-F’, Gcn5 expression is both in the nucleus and cytoplasm as this molecule shuttles between cytoplasm and nucleus. The staining pattern with the other Gal4 could be due to problems in the immunofluorescence protocol and acquisition parameters. We have now removed those images from Figure 1. Please check revised Figure 1.

      The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      We have used a pan-hemocyte promoter for the autophagy analysis to investigate if Gcn5 regulation over autophagy is a hemocyte specific effect which we indeed see. We have removed the western blot data now in the revised manuscript where we looked at Atg8 and p62 levels in whole larval lysates when Gcn5 was perturbed using hemocyte driver as the results were puzzling and difficult to comprehend given the complete absence of a p62 band in Gcn5 knockdown conditions. Also, it’s worth noting that Hml-Gal4 is also active in the LG hemocytes. We did 2 alternate promoters for prohemocytes to cross-validate some of our results and the chemical modulators experiment was done since effects like mTOR inhibition/nutrient sensing effects are systemic and hence such modalities were employed.

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      It is a possibility that Gcn5 perturbation could be affecting lymph gland size although we have not seen any consistent trend that would point towards this phenotype either upon knockdown or over-expression. We believe Gcn5 controls blood cell differentiation phenotypes strongly via mTORC1. But in order to answer reviewer’s comment we have now re-analyzed our crystal cell differentiation data particularly and quantitated it and represented it as crystal cell differentiation index for dome-Gal4 specific Gcn5 modulation and for the data with genetic modulation of mTORC1 pathway. Please see Fig 3P and S10J for the revised analysis.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      We thank the reviewer for their useful critique. We have now addressed this concern and we have genetically perturbed the mTORC1 pathway in the progenitors using both abrogation of TORC1 via depletion of Tor or Raptor or by activation using over-expression of Rheb. We have now included this data as Supplementary figures – Fig S10 and S11 and have described it in the results section. Please see results section “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The abstract could clearly be improved. It does not make a clear presentation of what is new in the manuscript. The conclusions that Gcn5 function in the lymph gland is mediated by the autophagy machinery and the acetylation of its non-histone target TFEB are not grounded and purely circumstantial. The implication of mTOR and nutrition in drosophila larval blood cell homeostasis has already been studied but not mentioned here. Most of the time the authors do not provide any possible explanation about the phenotypes they observe and how they fit with the current literature. Several pieces of results are of serious concern.

      We would like to thank the reviewer for their feedback. We have revised the abstract and have incorporated the insights obtained from our study. We have now included relevant literature that talks about the implication of mTOR and nutrition in Drosophila larval blood cell homeostasis (Please see Introduction section Paragraph 2 in the manuscript). We have also noted the input of the reviewer on many phenotypes lacking any description of a possible explanation. We have worked on results section to provide possible explanation and speculation wherever relevant.

      In the introduction, the authors do not provide an up-to-date and accurate presentation of the field. For example, they could use much more recent and comprehensive reviews since Evans et al. 2003. (eg. MID: 30733377 or 35887113). Their choice for signaling pathways involved in Drosophila blood cell progenitors seems very much biased for lead author self-citation rather than more directly related citations. It is surprising too that the authors failed to mention a series of publications on Akt/mTOR and nutrient sensing impact on drosophila larval blood cells (PMID: 22951642 ; 22911822 ; 22407365 ; 22510984). Along the same line, there are already several reports on autophagy genes implicated in Drosophila hematopoiesis and blood cell functions (PMID: 23406899; : 33560224 ; 20498061 ; 37623416). The introduction on GCN5 is a bit of a catalogue and should be streamlined- citing a recent review would be useful (PMID: 32735945). Again, the authors fail to cite publications showing that Gcn5 levels can be modulated by nutrition (PMID 27022023; 27874008) and they do not mention that amino acids starvation or mTOR inhibition leads to a decrease in GCN5 activity / TFEB acetylation (ref 40). Taking into account all the missing information, the novelty of the present manuscript is strongly decreased.

      We would like to thank the reviewer for the detailed suggestions on including the relevant literature that are appropriate and relevant to be mentioned in the context of the observations in our manuscript. We have now included these references and have cited them as per reviewer’s suggestions. Please check Introduction section paragraph 2.

      Results

      While there is little doubt that Gcn5 is expressed in the entire primary lobes based on Fig 1C-F, the quality of the staining in G-J (especially H, J) is really poor and essentially looks like non-specific background with no clear signal in the nuclei. Better images should be presented. The conclusion of the paragraph ("all cellular populations of the LG") and title of Fig.1 are not fully accurate as the authors do not provide evidence that Gcn5 is also expressed in posterior lobes.

      As per reviewer’s suggestions, since the images in Fig 1A-D’ clearly show that Gcn5 is expressed in the entire primary LG lobe in PSC cells, MZ and CZ; we have removed panels E-H’ which lacked clear nuclear signal. Fig1A-D’ clearly show the nuclear staining pattern of Gcn5. We have also modified the conclusion of the paragraph to say that Gcn5 is expressed in cellular populations of the primary lymph gland lobe accordingly.

      Concerning, Fig S1 and Fig 2, while the analysis seems technically sound, the results are puzzling. The lack of P1 differentiation in gcn5 null heterozygotes is very surprising. The authors should check that this stock does not carry a mutation in nimC1 (for details see: PMID: 23899817) and use other plasmatocyte differentiation markers to confirm their observation (also with the different allelic combinations). I'm also concerned by the levels of plasmatocyte differentiation and crystal cell number in the control line (notably in S1H), which seem very low (and quite variable for P1 as there is a notable difference between S1H and Fig 2H). Moreover, the analysis of the allelic combinations gives rather incoherent results: PCSC cell numbers are affected only in null/hypomorph, whereas differentiation (NimC1 and Hnt), as well as DNA damage, was only increased in hypomorph homozygotes. The authors propose no hypothesis to explain these observations.

      We have now repeated these experiments with the E333st null allele by placing it on a different balancer and we observe homozygotes that are alive till late third instar/early pupal stage as shown before by Carre et al., 2005. We have now included these revised results on the plasmatocyte differentiation status of the E333st heterozygotes and homozygotes (See Fig 2 and Fig S1). We do find P1 positive cells in the E333St heterozygotes unlike earlier. Plasmatocyte and crystal cell numbers in the control line always shows some level of heterogeneity. We have included the wild type control individually with those respective mutants during the experiment hence drawing a cross comparison across two different experiments would not be appropriate. We have now explained the observations obtained on PSC cell numbers (Discussion section paragraph 1). Experiments to check all hematopoietic aspects of the gcn5 null have been done after changing the balancer line and the null mutants overall show a decrease in PSC size and a widespread increase in hemocyte differentiation which could be due to a systemic effect due to various signalling pathways being affected which needs to be investigated and is beyond the scope of this study. This has also been discussed in the Discussion section Paragraph 1.

      Although a side-by-side comparison would have been better suited, it seems that the homozygotes or trans-heterozygotes do not have stronger phenotypes than the heterozygotes as far as crystal cell and DNA damage are concerned, which is rather unexpected. Besides the authors should introduce why they look at DNA damage.

      We agree with the reviewer that for the crystal cell and DNA damage phenotype the homozygotes or trans-heterozygotes do not have a stronger phenotype as compared to the heterozygotes alone but since these are whole animal mutants there could activation/inactivation of various signalling pathways and systemic effects that would be difficult to account for and comprehend here which needs to be investigated further. The only conclusion that we draw from these observations is that Gcn5 is required for maintaining blood cell homeostasis. Regarding DNA damage, we have now included the rationale and supporting literature for why we have studied DNA damage in the context of Gcn5. Please see result section 2 paragraph 1.

      Importantly too, the authors failed to obtain gcn5 E333st/E333st (null) larvae, whereas Carre et al. originally reported that E333st/E333st individuals are viable until the late third instar larvae. I suspect that the stock they use carries additional mutations that need to be eliminated by back-crossing it to control flies for several generations. Of note too, a recent report showed that a deletion of gcn5 (generated by CRISPR) does not prevent adult emergence, challenging the conclusion that gcn5 expression is absolutely required for fly development (PMID: 37545086).

      The reviewer is right in pointing out that E333st homozygotes survive until late third instar as reported by Carre et al.,2005. We have procured the null allele again and used another balancer to obtain homozygotes and we were able to get homozygotes that survived till late third instar as reported earlier. We have now included new data from these homozygotes for all hematopoietic aspects and heterozygotes particularly for plasmatocyte differentiation Please see Fig 2 and Fig S1 and corresponding results section 2 of the manuscript.

      Concerning the validation of Gcn5 knock-down and overexpression: the results are highly dubious. In Fig S2B (hml>Gcn5 RNAi), there is virtually no Gcn5 signal in the primary lobes but hml is normally expressed only in the cortical zone. How is it possible? Similarly, the western blot (which is really too much cropped around the bands of interest) does not show any signal in the hml>Gcn5 RNAi lane (not even some background. According to the Methods section, the western was performed on whole larvae extracts; hml-mediated knock-down can not wipe out its expression in all the tissues. As for the overexpression, flag immunostaining in hml>Gcn5-flag is mostly cytoplasmic (S2E), which doesn't make sense and does not fit with S2C (Gcn5 immunostaining).

      Hml-Gal4 is a pan hemocyte driver and its expression is not limited to the CZ of the primary lymph gland lobe (Banerjee et al., 2019) and recent single cell sequencing data corroborate this that Hml domain is not limited to the cortical zone (Yarikipati and Bergmann, 2026). GFP driven by Hml-Gal4 is spread out across the primary LG lobe which could explain the phenotype of no Gcn5 signal obtained in the immunofluorescence experiment. Regarding the western blotting experiment which was performed on whole larval extracts, we were also puzzled by lack of Gcn5 bands in these lysates upon depleting Gcn5 using Hml-Gal4. We need to systematically probe further to understand expression of Gcn5 in other tissues and organs. We have now removed the western blot data as the data obtained cannot be comprehended at the moment. Regarding the FLAG staining experiment – the staining gave us a cytoplasmic pattern and since Gcn5 is known to shuttle between the cytoplasm and nucleus it is possible that the anti-FLAG staining detected the Gcn5 localizing in the cytoplasm. It is difficult to draw a direct comparison here between the images S2C and S2E as both are different antibodies.

      The initial analysis of Gcn5 level modulation in the prohemocytes, PSC or Hml+ cells is mainly descriptive and the authors do not elaborate on possible explanations based on the current literature.

      We have added a possible explanation wherever required for these respective results on Gcn5 modulation in prohemocytes, PSC and Hml positive hemocytes. Please see result section 3 where we elaborate on possible explanation for the phenotypes observed.

      The structure/function analysis of Gcn5 is based on overexpression of truncated mutants in the prohemocytes using the tep4-GAL4 driver and monitoring PSC cell, prohemocyte maintenance, plasmatocyte and crystal cell differentiation as well as DNA damage. As the overexpression of the full-length protein was made with a different driver (Dome), it is difficult to interpret the data. Nevertheless, no clear message emerges from this analysis and the authors do not reach any conclusion. Thus, the interest of these experiments remains limited.

      The structure-function analysis was largely done to understand which of the domains of Gcn5 upon over-expression results in a dominant negative like phenotype and our analysis shows that expression of some of these domain mutants results in a dominant negative phenotype in the wild type genetic background which we have now stressed upon in the text. However, further mechanistic understanding and in-depth analysis of each of these domains of Gcn5 warrants further separate investigation and is beyond the scope of this study. Please see the end of result section 4 for conclusion and possible explanation.

      The authors then analyze autophagy markers (in hml>Gcn5 LOF or GOF). Contrary to their say, hml-GAL4 is not a pan-hemocyte marker. It would have been interesting to ensure that the effects observed on Atg8 and Ref(2)P in the lymph gland are cell-autonomous- as expected for a direct role of Gcn5 on this pathway. Again, it is very surprising that p62 is not detected in the western blot on whole larval extracts when Gcn5 is knocked down in Hml+ cells only (Fig 5D). Moreover, quantifications on multiple samples will be needed to validate the increase/decrease of p62 and Atg8 as detected by western blot. As for the RT-qPCR (Fig S5), according to the Methods sections, they were made on adult blood cells but this is not explicit in the result section.

      We have corrected the text and mentioned Hml-Gal4 as a hemocyte specific Gal4 shown earlier as Gal4 marking both embryonic and larval hemocyte population (Goto et al., 2003, Yarikipati and Bergmann, 2026). Regarding the Atg8 and Ref (2)P blots – yes, it is surprising to us too that the p62 is not detected in the larval lysates when Gcn5 is depleted using Hml-Gal4. However, this result was consistent over the replicates performed and needs to be further studied. Since this phenotype of complete absence of p62 in larval lysates upon Gcn5 depletion cannot be comprehended and explained, we have removed the western blot data from the figure and have just retained the immunofluorescence data and have also quantified the Atg8 and p62 puncta per cell and included this data in Figure 5, Graphs D and E. For the qRT-PCR we have now included a description in the corresponding results section. Please see result section – result 5 under “Autophagic flux in the Drosophila blood cells is negatively regulated by Gcn5”.

      The knock-down of TFEB or several autophagy genes in the prohemocytes (tep4-GAL4) leads to a rather convincing increase in plasmatocyte and crystal cell differentiation. It would have been interesting though to quantify prohemocyte maintenance, PSC cell number, and DNA damage. Also, the authors should have performed Gcn5 GOF/LOF experiments with the same driver (they present tep>Gcn5 RNAi in Fig 8 but without the proper controls).

      We have now included data for prohemocyte index (Figure S8M) upon knockdown of TFEB and other autophagy genes along with PSC cell number, DNA damage (Supple Fig S8) in the revised manuscript. Please see corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe” for the description of the results.

      The use of chloroquine should be better described. How long was the treatment? Did the authors observe an effect on autophagy in the lymph gland? Chrorloquine also affects lysosomal pH, so it remains to be demonstrated that the effects observed here are only autophagy-related.

      We have now written a detailed protocol for the treatment in the methods section and also mentioned the treatment time which is 16 hours in the results. We have included data to validate the effect of Chloroquine on autophagy by p62 and Atg8 staining in the LG and have quantitated the data (Refer Supple Fig S9) and the corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe”

      Similarly, the use of drugs to activate (3BDO) or inhibit (Rapamycin) mTOR should be better controlled. More generally, given the promiscuous roles of mTOR (and autophagy) in the larvae, tissue-specific manipulations would be better suited.

      We have now perturbed mTOR pathway genetically by activation and in-activation and have studied the effect on blood cell differentiation. Please see Figure S10 and the corresponding result section titled “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” where we discuss the results of genetic perturbation of mTOR pathway.

      Actually, as pointed out above, it has already been shown that modulation of Akt/TOR in hemocytes or amino-acid deprivation affects blood cell homeostasis (see above). The authors should definitely discuss how their results fit with the literature on this subject.

      We have added relevant literature in the introduction section and have also discussed how Gcn5 could fit into this context of nutritional sensing and control of hematopoiesis. Please check revised Introduction section paragraph 2. Also, check discussion section in last paragraph where we have discussed role of Gcn5 in nutrient sensing.

      Again, Gcn5 levels need to be quantified using multiple samples (Fig 7M, N) before concluding.

      Sorry for not including the quantitation earlier but we have now included the quantitation for the blots presented in Fig. 7 M and N.

      Finally, the authors show that 3BDO still induces an increase in blood cell differentiation when gcn5 is knocked-down in tep4+ cells and that Rapamycin still represses differentiation when Gcn5 is overexpressed in Dome+ cells. They conclude that mTORC1 overrides the effect of Gcn5. This seems a far-reaching conclusion given the available evidence.

      We have now toned down the conclusion that we make to accommodate other possibilities which we have been unable to test here currently.

      In particular, in the conditions used, the authors do not necessarily assess the activity/requirement for Gcn5 and mTORC1 in the same cell population.

      Other comments and suggestions:

      The discovery of the SAGA complex is not Grant 1999 but 1997 (PMID: 9224714).

      Ref 30 is not appropriate -nothing to do with HAT.

      GCN5 not only acetylates TFEB but also Atg7 (PMID: 28594263) to limit autophagy.

      Thank you so much for these suggestions. We have made the necessary amendments in the references.

      In the results section, the first paragraph is largely a repetition of the introduction. The same is true for most paragraphs in this section. A shorter (hypothesis-driven) introductory sentence would be more adequate.

      We have now taken the suggestion into consideration and made the necessary change in the results section throughout the manuscript.

      Fig 1: it seems that there is a higher accumulation of Gcn5 in a few cells in the cortical zone. This may correspond to crystal cells and could be easily confirmed.

      We have now checked this aspect. Please see supple fig S5 where we co-stain lozenge-GFP cells containing LG with Gcn5 to check for the accumulation. However, we do not see any accumulation in the Lozenge-positive crystal cells.

      Figure 3: the authors should also quantify the proportion of progenitors (dome>GFP+) in the different conditions.

      We have now done this and added it to the Figure. Please see panel N in Figure 3 and Figure S8M.

      Figure S3: how do the authors explain that Gcn5 knockdown in the PSC reduces plasmatocytes differentiation (but does not affect PSC cell number or crystal cell differentiation)? What could be the origin of the increase in DNA damage (essentially in CZ)? How do they explain that Gcn5 over-expression increases PSC size but does not affect (reduce?) blood cell differentiation?

      These observations need to be investigated further. We currently have no answer to these comments. The signals that are produced by the PSC could be affected due to which we observe these phenotypes like an effect on plasmatocyte differentiation and an increase in DNA damage whereas no effect on PSC cell numbers or crystal cell numbers which needs to be studied further. Also, in the case of Gcn5 over-expression in PSC we do not know how the increased size of PSC controls differentiation. This would need further experimentation and since this paper is not about the role of Gcn5 in PSC exclusively, we will look into this in our future studies. These aspects will be studied in our future follow-up studies as it is beyond the scope of the current manuscript.

      Figure S4: how do the authors explain the non-cell autonomous increase in PSC cell number upon Gcn5 KD/GOF in hml+ cells? How do they explain the increase in crystal cell number in Gcn5 GOF? Is it really cell-autonomous (i.e. all the Hnt+ cells are Hml+?)?

      We have discussed how Gcn5 depletion or over-expression in HmlΔ cells could affect PSC cell numbers. Please see discussion section, paragraph 1. Regarding the crystal cell phenotype - We have now tested if the increase in crystal cell numbers is cell autonomous by driving Gcn5 over-expression using a crystal cell specific driver and we find that the increase is cell-autonomous. Please refer to Supple Fig S5.

      The discussion is lengthy and should be reduced. It does not appropriately consider the current literature.

      We have tried to reduce the length of the discussion and have also added relevant references as per recommendations of the reviewer.

      Reviewer #2 (Recommendations For The Authors):

      (1) In general, it is not clear why in some of the experiments Tep-Gal4 is used to modulate proteins in prohemocytes while in others Dome-Gal4 is used.

      There is no particular reason. These Gal4’s have been used interchangeably as both label the hematopoietic progenitor population. Although recent single cell sequencing data has identified subsets within the progenitors namely core progenitors marked by tep4 largely and dome being a distal progenitor marker (Cho et al.,2020, Girard et al.,2021), in our study perturbations in Gcn5 using either of the Gal4’s results in a similar phenotype.

      (2) Considering alteration in lymph gland size (Figure 2), the number of positive cells should be analysed in relation to total cell numbers or s4ize.

      Although we do not find any visible differences or defects in the overall LG size in various genetic conditions discussed in this manuscript, we have done so for the plasmatocyte differentiation where we have represented it as plasmatocyte differentiation index (relative to the size of primary LG lobe) throughout the manuscript. We have now done this for crystal cell numbers too for critical genotypes in this manuscript and have represented it is as crystal cell index for example please see Figure 2O, 3P, S5G, S10J where these graphs have now been added.

      (3) Figure 1A G-I' does not look like mCD8 GFP expression, but rather cytoplasmic GFP.

      We have made the change in the figure and the corresponding text accordingly.

      (4) One of the main conclusions in the manuscript is that Gcn5 affects autophagy (Figure 5). Here, the puncta need to be quantified (relative to total cell numbers).

      Thank you for the suggestion. We have now quantitated the p62 and Atg8 positive puncta per cell and have represented it as panel D and E in Figure 5.

      (5) Figure 5 D and E show p62 and Atg8 total protein levels in the larvae when Gcn5 is modulated only in the hemocytes. It is surprising that there is a complete reduction in p62 levels across the whole larvae when Hml gal4 is used for the knockdown.

      Yes, we observe a complete absence of p62 in whole larval lysates when Gcn5 is depleted using Hml-Gal4 and we see this across replicates. This result is indeed puzzling to us and difficult to comprehend as to why a hemocyte specific driver would result in such a dramatic change hence we have decided to remove the western blot data as it is difficult to draw a solid conclusion from. We have retained the immunofluorescence data which shows a consistent alteration in autophagy upon Gcn5 perturbation using Hml-Gal4 and we have now included the quantification for the number of p62 and Atg8 positive puncta per cell for the IF data.

      (6) The beta-actin levels in the western blots in Figure 5 are highly oversaturated and do not represent loading control adequately. Also, it looks like there is substantially more total protein in 5D 3rd lane where Gcn5 is overexpressed.

      Thank you for pointing this out. We have loaded equal amount of protein in all the wells so we are unsure why the actin bands look over-saturated. We have now removed the western blot data from this figure as the data is puzzling and difficult to comprehend given a total absence of p62 in whole larval lysates in Gcn5 depletion conditions using Hml-Gal4. Hence, we are just retaining the immunofluorescence data.

    1. eLife Assessment

      This study uses convincing modeling methods and analyses of rich behavioral datasets to investigate the role of attention in value-based decision making; for instance, as when choosing between two snacks. The results are important, as they challenge existing theories that assume that paying attention to an available option biases the eventual choice toward that option. The results suggest that the correlation between attention and decision-making is formed largely after rather than before the (internal) choice process has terminated, a finding that offers an intuitively appealing rethinking of how attention and decision-making processes interact during value-based choices.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the weaknesses raised in the previous round of review.]

      Summary:

      This study examines whether gaze direction actively shapes choice during food preference decisions or whether gaze and choice evolve largely independently until the moment of commitment. The established framework in this context, the aDDM, assumes that gaze causally biases the accumulation of evidence in favour of the fixated item. The authors show convincingly that this model fails to fit key behavioural patterns across several datasets, as do other published models that make the same assumption. The authors propose an alternative model (Post-Decision-Gaze or PDG) in which gaze and decision formation are decoupled: gaze does not influence the decision process, nor is it drawn toward the ultimately chosen item, until after the decision threshold is reached. Only during the motor execution period (after commitment) is gaze directed to the chosen option. They demonstrate that this model fits several observed patterns better than the aDDM and related variants.

      Strengths:

      The work thoroughly considers multiple models and datasets. It advances an interesting alternative perspective on gaze-decision interactions and highlights meaningful shortcomings in existing models. The authors take the time to explain how modelling assumptions produce specific patterns in the data, which is certainly insightful to readers interested in the modelling of value-based decision making.

      Weaknesses:

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

    3. Reviewer #2 (Public review):

      Summary:

      Zylberberg et al. reanalyze eye-tracking and behavioral data to test two predictions of the attentional Drift Diffusion Model, finding that these predictions are not met. Similarly, predictions of normative models (inspired by rational inattention) are not in line with the data, and the authors propose a post-choice model of attention. This model better accounts for the two effects but also does not account for all patterns, so the authors conclude that eye movements most likely reflect both pre- and post-decisional processes.

      Strengths:

      A clear strength is the systematic falsification-based approach of the paper, establishing (partially) new predictions and testing to what extent these are met by extant models and by a newly developed theory. The authors do a good job in providing intuitions behind the effects and the reasons why models such as the aDDM predict them. The paper is of substantial relevance for the field, as it shows that effects pertaining to the last fixation(s) should be interpreted with caution. Another strength is the paper's transparency as the authors clearly acknowledge that their new model does not do a perfect job either.

      Weaknesses:

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

    4. Reviewer #3 (Public review):

      Summary:

      In this study, the authors reanalyzed choice, RT and gaze datasets collected from human subjects performing a food-choice task. They show that models that posit a causal role for attention in shaping the decision-making process fail to account for empirical observations in the data. These include the attentional drift diffusion model (aDDM) and models that derive attention-choice associations from an optimal policy. The authors show that a model that assumes that gazes are directed towards the chosen option after decision commitment captures more (but not all) empirical findings, suggesting that attention may reflect decisions once they are made instead of contributing to their formation. However, this post-decision-gaze (PDG) model failed to capture all aspects of the data, suggesting that gaze may reflect both decisional and post-decisional operations, and existing models are still missing some features of the gaze-directing process. The authors provide convincing evidence that post-decision gaze explains a number of empirical findings in this task.

      Strengths:

      (1) The analyses are generally appropriate, and the conclusions are supported by the data.

      (2) The study was rigorous, as the authors considered a number of alternative possible models for behavior, and evaluated their performance based on a wide range of qualitative predictions (as opposed to exclusively relying on model comparison).

      (3) The proposal that gaze may largely reflect post-decisional processes is interesting, and as far as I am aware, novel.

      Weaknesses:

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

      In particular, the "saccadic execution" parameter appears far too long and too variable to reflect merely execution; instead, it likely includes decisional components. This would make more sense since manual and saccadic planning essentially rely on distinct brain areas, hence it seems unrealistic that crossing a single threshold would trigger both manual and saccadic execution. Similarly, recovered manual non-decision times are substantially longer (though not more variable) than expected motor execution durations for button presses. These patterns suggest that parts of what the model treats as non-decision time are likely decisional in nature, although perhaps related to "action decision" rather than the "value-based decision" of interest to the authors. To what extent these two processes neatly follow each other or overlap could be usefully considered.

      We have added a paragraph to the Discussion explaining how our model’s estimates of sensory and motor latencies relate to corresponding values inferred from physiology or behavioral manipulations (e.g., Bompas et al., 2024). Specifically, we write:

      “The key assumption of the PDG model is that there is a delay between the moment a choice is internally committed and the moment it is externally reported with a key press. Because eye movements are typically faster than manual responses (𝜏<sub>e</sub> < 𝜏<sub>m</sub> in our simulations), this delay creates a window during which gaze can already be directed toward the covertly chosen item before the response is formally registered. We do not interpret these non-decision latencies as irreducible physiological minima for moving the eyes or pressing a button (Bompas et al., 2025). Rather, they are inferred indirectly by fitting an additive non-decision-time parameter to the behavioral data, which we decompose into a sensory delay (𝜏<sub>s</sub>) and a manual execution delay (𝜏<sub>m</sub>). Values of 𝜏<sub>e</sub> are then chosen so that the model reproduces the observed magnitude of the behavioral effects. This estimation procedure has important limitations. Some participants show relatively “flat” chronometric functions: response times vary little with value despite otherwise normal psychometric performance. Such patterns likely reflect processes not explicitly represented in the model, including procrastination, reduced motivation, task-unrelated thought, or noise in item ratings. Within a drift-diffusion framework, however, these cases are accommodated by assigning a long non-decision time together with a short evidence-accumulation period (Table S1). Consequently, some estimated non-decision times are substantially longer than would be expected if they represented only sensory and motor delays. A further limitation is conceptual. We model non-decision time as occurring either before or after evidence accumulation, whereas in reality decisional and non-decisional components are likely temporally interleaved (Graziano et al., 2011). This simplification may also inflate the recovered latency estimates. With these caveats in mind, sensory and oculomotor delays on the order of 300 ms remain broadly plausible, although they likely lie near the upper end of a realistic range. The estimated eye-movement latency is especially long. For instance, in monkeys trained to report simple perceptual decisions with a saccade, roughly 100 ms elapses between the threshold-crossing signal in parietal cortex (or the superior colliculus) and the executed eye movement (Roitman and Shadlen, 2002; Stine et al., 2023). Crucially, however, varying the assumed non-decision latencies across a reasonable range does not alter the qualitative predictions of the model (Fig. 8).”

      Further, we have added a parameter sensitivity analysis. Importantly, although the magnitude of the predicted effects depend on the non-decision latencies, the qualitative aspect of these predictions do not (new Figure 8). Specifically, (i) the increasing tendency to look at the ultimately chosen item as time elapses (new Fig. 8A), (ii) the lack of an interaction between the last-fixation bias and overall value (Fig. 8B), and (iii) the absence of an effect of choice consistency on Δdwell (Fig. 8C) are all findings that are independent of 𝜏<sub>e</sub>.

      Reviewer #2 (Public review):

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Following this suggestion (and the reviewer’s elaboration in the private comments to the authors), we have substantially restructured the manuscript. Both aDDM predictions are now presented together (new Fig. 2), and Figs. 3–4 test these predictions across multiple food-choice datasets. In doing so, we no longer treat the data from Krajbich et al. (2010) separately, and we extend the analysis of the last-fixation–choice association (MELFB) to additional datasets. We note that the same datasets could not be used in both Figs. 3 and 4, as some lack information on the final fixation required for the MELFB analysis. Nevertheless, results are highly consistent across datasets and align with findings from a recent study by Ting & Gluth (2025), which independently identified and examined one of our key predictions; this work is now cited in the revised manuscript. Finally, to reduce redundancy, we have consolidated all aDDM variants and optimal models into a single figure (new Fig. 10).

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

      The key predictions of the model depend on the difference between the manual (𝜏<sub>m</sub>) and eye-movement-related (𝜏<sub>e</sub>) latencies. We have now added a parameter-sensitivity analysis to show how the model predictions depend on this difference. The new analysis shows that while the quantitative predictions do depend on the precise latency values, the results are qualitatively similar across values of 𝜏<sub>e</sub> (new Figure 8).

      Reviewer #3 (Public review):

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

      Thank you for this suggestion. We added a new paragraph to the discussion (paragraph #2), where we argue that it is sensible for a decision maker to direct the gaze to the chosen item once a covert choice commitment has been made, as the benefits of attending to a stimulus do not end with the decision itself. Specifically we now write:

      “Instead, these observations are better explained by a post-decision account of the gaze-choice association that is, one in which gaze shifts to the selected item after a covert commitment to a choice. We argue that directing gaze to the chosen item after a covert choice commitment is sensible, as the benefits of attending to a stimulus do not end with the decision itself. In naturalistic settings, for instance, selecting a food item is typically followed by the action of reaching toward it, where visual attention supports spatial localization and motor planning for the upcoming action. Although participants in our computerized task did not physically act on their choices, these sensorimotor processes are likely highly automatized and may still be engaged by default, even when not strictly required. Beyond motor preparation, post-decisional attention may also serve additional functions, such as facilitating sensory anticipation of the reward, supporting metacognitive evaluation of the decision, and contributing to value updating for future choices. From this perspective, a degree of attentional “stickiness” whereby the chosen item remains preferentially attended after commitment could emerge as an effectively optimal policy once these post-decisional processes are taken into account. Moreover, a specific feature of the task design may further reinforce this tendency: in the snacks paradigm, the unchosen item typically disappears from the screen immediately after a response is registered. It is therefore plausible that directing gaze to the chosen item after commitment partly reflects anticipation of the imminent disappearance of the unchosen option. To disentangle these mechanisms, it would be interesting for future work to test whether this attentional bias persists when the chosen item, rather than the unchosen one, is the stimulus that disappears upon response.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments:

      (1) Framing of the modelling approach

      The manuscript would benefit from acknowledging the known limitations of DDM-based frameworks, especially given that the entire study is conducted within these constraints. The introduction highlights successes of the DDM, but the manuscript does not mention any of its conceptual or empirical limitations.

      We are unsure about what specific limitations the reviewer has in mind, but we have added a paragraph to discussion mentioning some limitations, like the inflation of the non-decision times and the difficulty of interpreting the fit parameters (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (2) Dependence on non-decision time assumptions

      The alternative model's explanatory power appears to rely heavily on assumptions regarding the decomposition of non-decision time: fixed visual encoding (𝜏<sub>s</sub>= 0.3 s), manual non-decision time (𝜏<sub>m</sub>; two free parameters), and saccadic execution (𝜏<sub>e</sub>; fixed parameters μ<sub>e</sub> = 0.35, σ<sub>e</sub> = 0.11).

      - 𝜏<sub>e</sub> is substantially longer and more variable than typical saccadic execution times, suggesting it likely incorporates decisional components.

      - Estimated 𝜏<sub>m</sub> values are approximately twice as long as known manual execution durations.

      - σnd is more plausible, implying that variability is captured correctly but mean durations are not.

      Together, these points raise the possibility that portions of what the model treats as non-decision time are in fact part of a (action) decision process. Only then does it make sense to assume that Tm is usually larger than Te. If Tm and Te were truly execution delays, then Tm would always be larger than Te.

      You may find it helpful to consider the framework in Bompas et al. Psych Review (2024), which discusses in detail what non-decision time is likely to comprise across effectors.

      Thank you we have added (i) a sensitivity analysis showing that our results are robust to changes in the specific value used for the eye movement related latencies (new Fig. 8), and (ii) a new paragraph in Discussion addressing the issue of the mismatch between our parameter estimates and the manual and saccadic execution times (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (3) Code availability.

      The authors should consider sharing all relevant code and data publicly.

      We agree, we now share the code and data on GitHub and indicate so in the revised manuscript.

      Minor Comments:

      (1) Lines 74-77. These are not worded as predictions but as questions; one tests predictions, but answers questions. I feel it would be clearer to stick to predictions (like in the abstract), and the introduction could benefit from explaining these predictions in a bit more detail (I found it difficult to get my head around these predictions from the intro text only).

      We rewrote the section in the introduction where we provide a gist of the model predictions (last paragraph of Introduction). We agree with the reviewer that the previous explanation was not clear.

      (2) It is confusing that panel B appears to the left of panel A in Figure 2.

      We agree. We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed and the panels follow a more logical order.

      (3) Figure 3C - remove MATLAB toggles.

      Yes, thanks.

      (4) Figure 5A shows the proportion of left choices, but the text and legend refer to right choices.

      Good catch, thank you.

      Reviewer #2 (Recommendations for the authors):

      This may appear self-serving, but the authors seem to be unaware of some highly relevant work from our group. Most importantly, in a recent publication (Ting & Gluth, 2024, JEP General), we have already looked at the dependency of the last- (or final-) fixation bias on overall value in value-based (VB) and perceptual (P) decisions. In VB, we found a negative effect; in P we did not find a significant effect. This is largely consistent with the current results, showing a negative but not significant trend. Another relevant work is Gluth et al. (2020, Nat Hum Behav), where we extended the aDDM by assuming that the probability to fixate on an option is a function of the accumulated evidence for that option. It would be interesting to know whether this assumption changes the predictions of the aDDM. Finally, we just published a new theory on how people search for information to make efficient value-based decisions (Gluth et al., in press, Psychol Rev; https://osf.io/preprints/psyarxiv/3qzak_v2). Although this theory focuses on multi-attribute choices, it can be applied to "simple" choices, too (by assuming that there is only one attribute = value). Interestingly, while the model also mispredicts a (slight) increase of the last-fixation bias with overall value, it correctly predicts the independency of the dwell-time advantage effect on choice consistency as well as the small increase of the effect with RT (attached here is a figure to show this: [https://elife-rp.msubmit.net/elife-rp_files/2026/01/22/00149589/00/149589_0_attach_9_477122. pdf], and the match with the empirical data shown in Figure 3B and 12 is striking). In general, the model shares many features of the Callaway and Jang models, but does not need to assume a biased value prior, which the authors suggest is responsible for the misprediction of the second effect. I leave it up to the authors to discuss this new theory, but I wanted to point this out.

      Thank you for pointing this out; these are all relevant points and studies.

      We now note that the first of our predictions has recently been identified and tested by Ting and Gluth (2025).

      We also considered extending the manuscript with a variant of the model proposed by Gluth et al. (Psychological Review, 2026). In fact, we attempted to fit this model to the Krajbich et al. (2010) dataset under the assumption that the duration of each sampling epoch is a free parameter. We find this model very interesting. However, in our current implementation it appears to make the same qualitative prediction as the aDDM, namely that ΔDwell depends on choice consistency (see Author response image 1).

      Given this, we have decided not to include these results in the manuscript. It remains possible that with further development particularly with a more realistic specification of fixation durations (e.g., allowing them to depend on value) the model could account for the full set of observed effects. We think this would be best addressed in a separate study.

      That said, we do find the model promising, as it provides a better account than most of the alternative models we explored for the patterns shown in panels D, H, and I.

      Author response image 1.

      Fits of a variant of the MACS model (Gluth et al. 2026) to the data of Krajbich et al. (2010).

      The paper would benefit substantially from restructuring. The aDDM's predictions are provided first, together with the empirical data, and then the optimal models are discussed. But Figure 2 shows all of this together. Later, the new (PDG) model is elaborated, and its predictions are shown. Towards the end of the results, variations of the aDDM and combinations of aDDM and PDG are shown in a series of figures (8-11), followed by a last figure showing one of the tested effects in other datasets. All of this feels pretty much thrown together without a clear structure. For instance, the aDDM and the optimal models could be described together (or the optimal models get a separate figure). The additive variants could be described earlier. And some figures could be put into the supplement. And the empirical results of the different studies could be shown together.

      We fully agree with this suggestion. We have now restructured the manuscript along the lines proposed by the reviewer (see the more detailed explanation of the restructuring in our response to the public comments).

      I strongly suggest avoiding the term "influence" in the y-axis of Figure 2, upper row, as it implies causality. Similarly, in line 182, the term "causal influence" is used in the context of the Callaway model, but as far as I know, this is not what the model assumes.

      We replaced the y-axis label with “Association of last dwell with choice (β)”

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 2 - Panel labels for A and B are reversed?

      We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed.

      (2) Does 3C include a .pdf screenshot?

      Thank you, it’s a Matlab bug on Mac. I guess they want us to switch to Python -:)

      (3) Figure 4 - It would be helpful if the green line were defined in the figure legend.

      Added

      (4) The effect size in 5B looks much more dramatic than in 2B(A?) - Is this for one example subject as opposed to all subjects? Please clarify what is different about the data.

      We are no longer showing the psychometric functions in Figure 2.

      (5) Line 252 - they say they compared the probability of choosing the right item (Fig. 5B) by the y-labels of that figure, which are all p(choose left).

      Yes, corrected now.

      (6) In general, they reference the subpanels of Figure 5 out of order, which causes the reader to jump around. They might consider reordering the panels of the figure so they follow the ordering of descriptions in the text.

      We agree, we have rearranged the figure panels to follow the ordering of the descriptions in the text.

    1. eLife Assessment

      This important study measures single-unit activity in area MT of awake-behaving monkeys to test the idea that sensory adaptation contributes to flexible evidence accumulation during decision making. The authors provide compelling evidence that adaptation to different temporal contexts shapes both perceptual judgements and neural responses. Although the precise computational mechanisms underlying these effects remain uncertain, the results support the conclusion that recent sensory history influences the temporal dynamics of decision formation. This work will be of interest to researchers studying visual perception, sensory adaptation, and decision making.

    2. Reviewer #1 (Public review):

      McGaughey and Gold ask where in the decision process the flexibility of evidence accumulation arises, proposing that it is not solely a property of downstream integrators but is also supported by stimulus-specific sensory adaptation in the middle temporal area (MT). Recording single-unit activity in rhesus macaques during a motion direction-discrimination task in which an adapting stimulus of varying temporal stability precedes an identical test stimulus, they find that more rapidly changing contexts produce weaker and less discriminable MT responses to the test stimulus, which they argue accounts in part for context-dependent changes in decision-making behavior. Through session-level correlations they further identify pupil-linked arousal as a parallel, apparently separable contributor.

      The main strength is the shift of perspective toward the encoding stage: rather than treating MT as a static input to flexible downstream integrators, the authors show that early sensory cortex can itself contribute adaptive, context-dependent signals that shape behavior. The conceptual advance is supported by a well-designed paradigm-total exposure to each motion direction is matched across conditions and the test stimulus is held identical-together with single-unit recordings and simultaneous pupillometry. The behavioral effect is consistent across three animals, and the fact that context-dependent differences emerge over repeated stimulus presentations within a trial, rather than as a sustained baseline offset across blocks, ties the effect convincingly to stimulus-specific adaptation.

      The behavioral effect constrains the temporal dynamics of decision formation but does not uniquely identify its algorithmic basis: a leak, a saturating non-linearity, or a reduction in the gain of integration are all compatible with a shallower rise of accuracy with viewing time, and the reduced MT discriminability is itself an encoding-stage efficiency effect of this kind. The manuscript appropriately treats the algorithmic basis as unresolved, noting that distinguishing these accounts would require analyses not available here, such as reverse-correlation or motion-energy kernels with lower-coherence test stimuli.

      The inference that the adaptation- and arousal-related signals operate independently rests on the absence of session-wise correlations between the neural and pupil measures and their behavioral contributions. Given the noise in the trial-wise estimates, this is best read as consistent with, rather than demonstrating, true independence, as the authors note.

      Overall, the authors largely achieve their aim of showing that sensory adaptation in MT shapes the evidence available for time-dependent perceptual decisions. The evidence for a sensory-encoding contribution is convincing, while the claim of independence between adaptation and arousal is more tentative and is framed as such.

    3. Reviewer #2 (Public review):

      McGaughey and Gold trained rhesus macaque monkeys to perform a motion-direction discrimination task in which a behaviorally irrelevant adapting stimulus with either fast or slow direction alternations preceded a variable-duration test stimulus, while simultaneously recording single-unit activity in area MT and pupil diameter. They report that adaptation to the more rapidly changing stimulus was associated with reduced behavioral sensitivity, attenuated test-evoked MT responses, and larger pupil-linked arousal signals. The authors interpret these behavioral changes as evidence for context-dependent adjustments to the temporal dynamics of decision formation and argue that these adjustments are supported by both sensory adaptation in MT and arousal-related mechanisms. More broadly, they conclude that flexible evidence accumulation in dynamic environments arises from distributed adjustments across sensory encoding and neuromodulatory systems rather than solely from changes within a downstream accumulator. If correct, this interpretation has important implications not only for our understanding of perceptual decision making, but also for broader theories concerning the functional role of sensory adaptation.

      The conclusions of the paper are generally supported by the data. Evidence for adaptation-induced changes in sensory encoding, behavior, and pupil dynamics is convincing, and the revised manuscript substantially strengthens the connection between the behavioral findings and the proposed decision-making framework.

      Comments on revised version.

      The revised manuscript provides a clearer account of how recent stimulus history influences behavioral performance. In the original version, aspects of the psychometric functions were interpreted as evidence for a more leaky evidence-accumulation process, although some of these effects could potentially have reflected alternative mechanisms, including influences of the adapting stimulus on short-duration trials. The additional analyses and discussion included in the revision clarify that information from the adapting stimulus contributes to behavior at short viewing durations and appropriately temper claims regarding the specific computational mechanism underlying the observed behavioral effects. While the data do not uniquely identify whether these effects arise from changes in leak, other nonlinearities, or related decision processes, they provide convincing evidence that recent temporal context influences the temporal dynamics of decision formation.

      My original review also noted that different sections of the manuscript relied on different behavioral metrics and analytical approaches when relating behavioral changes to neural and pupil-linked measures. The revised manuscript now provides a clearer rationale for these choices, including distinctions arising from the different trial types and time windows used in the neural and pupil analyses.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) Alternative mechanisms for performance differences.

      The authors assume that the difference in performance between the low-switch (LS) and high-switch (HS) frequency conditions is explained by a change in the "leakiness" of integration. However, several other mechanisms could potentially explain this effect:

      (1) Temporal Uncertainty: Integration might start later in the HS condition, leading to lower performance.

      (2) Reduced Efficiency: Integration could be less efficient in the HS condition (i.e., lower signal-to-noise ratio) without a change in the leak parameter itself.

      (3) Evidence Contamination: Motion information from the adapting stimulus in the HS condition may be integrated rather than ignored, which might be the case since the transition from the adapting to the test stimulus is not externally cued.

      To distinguish between these alternatives, I suggest two possible analyses. First, a formal model comparison could be performed, though I acknowledge this may be inconclusive in the absence of response-time data. Second, an analysis of motion energy kernels could be revealing; the leak hypothesis makes the specific prediction that for long test stimuli, early samples should contribute more to the choice in the LS condition than in the HS condition, relative to late samples.

      We thank the reviewer for raising these important points. We agree that we cannot definitively identify the algorithmic underpinnings of the behavioral effects we report and have made substantial revisions to the manuscript to be clearer about what is supported and what is speculative in our claims. Most importantly, we agree that we do not know if the context-dependent differences in how accuracy depends on viewing time are based on adjustments to a leak or to something else (e.g., a saturating non-linearity, as we identified in Glaze et al, 2015, that is separate from the leak itself), which we cannot resolve with this dataset, even with more formal model comparisons. We therefore:

      Changed the wording throughout the manuscript to refer to changes in leakiness as just one of several possible sources of the behavioral differences. We also added this point to the list of “limitations” (and possible future directions, including using motion-energy kernels, which would require us to use lower-coherence test stimuli) in the Discussion (L487-493).

      Added a new figure panel (Fig. 2D), a new Extended Data figure (Extended Data Fig. 3), and additional explanatory text (L168-175) that collectively describe the behavior in more detail, including quantifying a “crossover” dynamic similar to what we reported previously (Glaze et al, 2015).

      Added new explanations (L152-163) and analyses (Extended Data Fig. 9) indicating that the monkeys used some information from the end of the adapting stimulus to inform their decisions, which accounts for the patterns of choices at the shortest viewing durations.

      Indicate that the context-dependent differences in the slopes of the psychometric functions (and complementary analyses based on “raw” accuracy measures as a function of binned viewing duration) rule out the temporal uncertainty and evidence contamination explanations, but are consistent with effects on the temporal dynamics of the decision process (L175-179).

      (2) Independence of neural and pupil-linked signals.

      The authors take the lack of session-wise correlation between context-dependent contributions from neural and pupil terms as evidence that these two signals provide independent contributions to the behavioral effect. However, could this lack of correlation simply be a result of high variability or noise in these estimates? The data shown in Figure 7B suggests that measurements are very noisy, which might obscure a potential relationship.

      We agree that the lack of session-wise correlation between neural and pupil terms cannot be taken as definitive evidence of independence. We have both softened the language around the claim (L368) and added a sentence to the Discussion (L464-468) acknowledging that this lack of correlation may reflect underlying noise and/or variability rather than true independence of the underlying mechanisms.

      Reviewer #1 (Recommendations for the authors):

      (3) The neural data analyses rely fundamentally on "switch" trials (Figures 3-5). It might be informative to also examine "non-switch" trials to see if there are specific neural markers indicating the exact moment the motion stimulus becomes behaviorally relevant. Given that this may fall outside the primary focus of the paper, it is up to the authors whether to pursue this line of inquiry.

      We thank the reviewer for this suggestion. We agree and have added new analyses of data from non-switch trials (Extended Data Fig. 9), which show some effects of stimulus information from the adapting epoch on the monkeys’ choices, as we detail below in response to related comments from the other reviewers.

      Reviewer #2 (Public review):

      Aspects of the behavioral analysis would benefit from a tighter connection between theoretical claims about evidence accumulation and the empirical features of the psychometric functions. For example, the rightward shifts observed across adapting conditions are interpreted as consistent with a reset of accumulation on switch trials, but similar patterns could also arise from failures to detect the test stimulus on a subset of trials, leading responses to default to the final adaptor direction. Likewise, changes in psychometric slope and asymptote are attributed to differences in evidence accumulation without explicit modelling or consideration of alternative explanations.

      Clarifying how specific features of the psychometric functions map onto distinct components of the decision process will strengthen the link between the theoretical framework and the behavioral data.

      We agree and have made substantial revisions to address these important points. Specifically, we added a new figure panel (Fig. 2D), new Extended Data Figures (3 and 9), and several lines of explanatory text (L152-179) that collectively describe the behavior in more detail, including clarifying that: 1) for the shortest viewing durations, the monkeys’ decisions were informed by information from the adapting stimulus, which accounts for generally lower accuracy on LSF (longer exposure to the final adapting direction, thus more accumulated evidence for that direction before processing the switch) vs. HSF (shorter exposure to the final adapting direction, thus less accumulated evidence for that direction before processing the switch) switch trials; and 2) as viewing duration increased, the rate of rise of accuracy versus viewing duration was higher for LSF vs. HSF trials, implying differences in the process of evidence accumulation. As detailed in our response to a similar comment from Reviewer 1, above, we are now careful to temper our claims about the specific computational basis (e.g., a leak or other form of nonlinearity) for these differences.

      We also de-emphasized our treatment of the asymptotes of the psychometric functions. In principle, these regimes could give insights into leakiness (which can limit the total amount of information that can be accumulated) and lapses (which are measured at the asymptotes). In practice, however, the long-duration trials that constitute the asymptotes were relatively under sampled (to promote the unpredictability of the offset of the stimulus, which we believed was the more important consideration when designing the experiment), yielding unreliable estimates.

      A slight concern is the lack of a consistent analytical approach for relating behavioral changes to neural and pupil-linked measures. Different sections of the manuscript rely on different behavioral metrics-such as differences in accuracy within a selected stimulus-duration range (e.g., Figure 5C) or psychometric slope differences (Figure 6C) without clear justification for these choices. The analytical approach likewise varies between simple correlational analyses (Figure 5C, Figure 6C), pseudo-experimental group comparisons (Figures 5D, E), and the inclusion of neural or pupil terms in the behavioral psychometric regression model (Figure 7B). While each metric and approach may be defensible in isolation, adopting a more consistent framework will help convince readers that the reported effects are robust and not contingent on the selective choice of metric or analysis.

      We thank the reviewer for this thoughtful critique and agree that the rationale for our choice of behavioral metrics and analytical approaches could be stated more clearly. We have added text to the relevant sections of the Results (L247-251) clarifying these choices. In particular:

      The neural analyses (Figures 3D-E, Figure 4, Figure 5D-E) focused on preferred-motion switch trials, because: 1) low switch-frequency non-switch trials provide an additional 800 ms of exposure to the final adapting-stimulus motion direction relative to high switch-frequency non-switch trials, which confounds comparisons of context-dependent evidence encoding between conditions, and 2) MT neurons exhibit minimal responses to null motion (although note that we also included analyses based on ROC area, which is computed from both preferred- and null-motion switch trials, to account for possible contributions of null-motion responses; Figure 5A-C). Thus, to ensure a meaningful comparison between neural and behavioral measures, we used behavioral accuracy on switch trials as the relevant metric in Figure 5C-E, rather than psychometric slope, which is estimated across both switch and non-switch trials.

      The pupil analyses (Figure 6) focused on a time window preceding test-stimulus onset, representing the arousal state around when the decision process started, and included both switch and non-switch trials. Thus, for these analyses we used psychometric slope, which is estimated across both switch and non-switch trials.

      We used several different analyses to compare and contrast the neural-behavioral and pupil-behavioral relationships because they provide complementary and useful insights. The correlational analyses in Figures 5C and 6C characterize session-level relationships between neural/pupil signals and behavior. The group comparisons in Figures 5D–E provide a complementary visualization of the same relationship. The model-based approach in Figure 7 then allows direct quantification of the trial-wise contributions of each signal to behavior within a common framework. Importantly, the conclusions drawn from each approach converge on the same interpretation, which we believe speaks to the robustness of the reported effects.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 2 legend. Description of 'running average (5-trial window)' is unclear - presumably this is a running average in stimulus space rather than across trials.

      We thank the reviewer for flagging this ambiguity. We have updated the legend (L136-137) to clarify that the running average is computed across trials sorted by test-stimulus duration.

      (2) L158. Difficult to establish an asymptotic performance level for HSF conditions within the stimulus duration range tested.

      We have removed the reference to asymptotic performance and replaced it with a discussion of performance on longer-duration switch trials in the context of the newly added Figure 2D.

      (3) L515 Equation 1. While this is a standard formulation of lapse rate in psychometric functions, the construction here in terms of switch probability is not standard. Given the task and training, it seems more likely that on lapse trials, the animal will respond according to the last adapted direction (rather than randomly switch/stay with equal probability).

      We thank the reviewer for this point. We agree that it is possible that on at least some of the “lapse” trials the monkeys may respond according to the final adapting-stimulus direction rather than choosing randomly. However, we cannot distinguish those alternatives using this task design. We include a statement to this effect in Methods (L569-571).

      To explore the idea further, we refit the behavioral data using separate upper and lower asymptotes corresponding to lapse rates on switch and non-switch trials, respectively. Across monkeys, there were no significant differences between upper and lower lapse rates for either low (Wilcoxon signed-rank test for equal medians: p = 0.15, Cohen's d = -0.13) or high switchfrequency (p = 0.07, Cohen's d = -0.16) conditions. So, at the very least, there was no evidence for lapse-like errors driven by switch- (or non-switch-) specific defaults to the final adapting direction.

      (4) L256. Statistical significance of attenuation is not directly tested here.

      We have replaced "were attenuated" with "we did not identify any reliable context-stability differences" (L297) to accurately reflect what was directly tested without implying a statistical comparison between groups of sessions that was not performed.

      (5) L429. Does the increase in explanatory power warrant the increased complexity of the model here?

      We thank the reviewer for raising this important point. We used Tjur's pseudo-R<sup>2</sup> because it does not increase by default with added model complexity, making it more conservative than other R<sup>2</sup> measures in this respect. Tjur's pseudo-R<sup>2</sup> is a coefficient of discrimination, and as such its value increases only when additional terms improve the model's ability to separate predicted probabilities across response outcomes. Thus, the observed increases in explanatory power when adding neural or pupil terms reflect real improvements in discriminability rather than an artifact of model complexity. We have added a brief clarification of this point to the Methods (L662-664).

      Reviewer #3 (Public review):

      The task design may not be optimal. While the amount of time the monkey is exposed to each motion direction during the adapting stimulus is matched, it's hard to know if the reduced MT responses to the test stimulus are truly due to the greater frequency of switches during the HSF adapting stimulus or because the monkeys have been exposed to more repetitions of the stimulus. It's increased sensory adaptation in either case, but it makes it problematic to interpret this as temporal context-dependent adaptation specifically. I think this could potentially be partially addressed by an analysis that is in the paper, but could potentially be emphasized/fleshed out more, specifically the results shown in Figure 4D that seem to show that most of the reduction in neural response for adapting units occurs between the first and second stimuli.

      The reviewer raises an important point. The number of stimulus repetitions and switch frequency are confounded in the experimental design, making it difficult to attribute context-dependent differences in MT responses to the temporal pattern of switches rather than to accumulated repetitions. We also note, as the reviewer acknowledges, the observed differences reflect sensory adaptation either way. Figure 4D does offer relevant evidence, suggesting that a majority of the change in neural response occurred with just one stimulus repetition. This finding complicates an interpretation where adaptation scales with the number of stimulus repetitions. We have added several lines to the Results about these points (L231-233).

      The pupillometric analysis seems to be an indirect way of assessing whether the accumulator itself might be modulated by temporal context, but the link could be made clearer. The authors show that context-dependent behavior is related to pupil size, which is related to arousal/neuromodulation, but it would be helpful to have some idea of what neural mechanisms underlying adaptive decision-making are actually impacted by this neuromodulation. Lacking neural data to address this question (e.g., from a brain region proposed to be involved in the accumulation process), at least more discussion of this would be helpful. Essentially, I'm unsure of how to interpret the pupil results: the argument that temporal context affects instantaneous evidence encoding in MT that then drives the accumulator is very clear, but I am a bit confused about what, mechanistically, I should think about the effect of neuromodulation doing.

      We thank the reviewer for this thoughtful comment and agree that the mechanistic interpretation of the pupil results could be made clearer. We acknowledge that we cannot directly identify the neural mechanisms underlying the arousal-related contributions to adaptive evidence accumulation from pupil data alone, given that pupil size is an indirect and imperfect proxy for neural (e.g., LC-NE system) activity. However, we can offer some informed conjecture and have added to the Discussion (L469-482) in an effort to elaborate on possible mechanisms.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract could be retooled - does not emphasize the pupillometry/arousal results very much, and they are presented more as a control than an independent result.

      We agree and have revised the Abstract accordingly.

      (2) Do all neural/pupil analyses use only switch trials? Sometimes the figure captions do specify only switch trials, but not everywhere. It would be helpful to specify either in the Methods or at the beginning of each figure caption that all subplots show switch trial results. Also, if you do always use switch trials, it would be useful to see in the Supplement how the non-switch trial results differ from switch trials. It seems like they may in interesting ways based on the behavioral results (supporting a reset of evidence accumulation on switch but not non-switch trials).

      We thank the reviewer for flagging these important points. We have added a justification for switch trials (L186-190) as well as clarification about which trial types were used for which analyses (L246-249) and information about trial types to relevant figure captions. We have also added a new Extended Data figure (Extended Data Fig. 9) examining relationships between neural activity and behavior on non-switch trials. As inferred by the reviewer, behavior on non-switch trials is consistent with the use of information from the adapting stimulus.

      (3) In Figure 3C, 5B, etc, when computing firing rate for the test stimulus (50-500 ms), are differently sized windows used to compute the rate for different test stimulus durations (since some will be <500 ms)? Or are only trials where the test stimulus duration is > 500 ms used for this analysis?

      We thank the reviewer for raising this point. To clarify, the 50–500 ms window does not reflect a fixed window applicable for all trials. Rather, neural activity from 50 ms after test-stimulus onset through test-stimulus offset was included for each trial, with 500 ms serving as the upper bound for trials with longer durations (> 500 ms). We have clarified this in the Methods (L607-610) to avoid ambiguity.

      (4) I think it might be better to be consistent with the time windows used for analysis; specifically, to choose either the 50-500 ms window used in Figures 3, 4, and 5B, or the 200- 400 ms window used for the remaining analyses in Figure 5.

      We agree that using the same window for all of the analyses would improve consistency, but not doing so provides advantages that we believe take precedent and now describe in more detail. The broader 50–500 ms window used for Figures 3, 4, and 5B was chosen to characterize MT neural activity over a relatively large a time window, ensuring that every trial contributes to each estimate. Because test-stimulus durations were drawn from a truncated exponential distribution (100–1200 ms), restricting these analyses to the 200–400 ms window would have excluded the substantial proportion of trials with durations <200 ms (but would yield similar figures and conclusions). The narrower window used in subsequent analyses allows us to focus on the conditions that exhibited the biggest modulations of neural activity when comparing them to behavior.

      (5) Similarly, provide justification for using only trials ending 375-600 ms after test stimulus onset for the behavioral correlations. It seems reasonable to choose a subset of test stimulus durations where the monkeys' behavior is greater than chance but less than ceiling, but it would be good to specify this so that it doesn't seem arbitrary.

      We agree and have added text to make this important point (L249-251).

    1. eLife Assessment

      By investigating spine nanostructure and dynamics across multiple genetic mouse models for neurodevelopmental disorders, this important study has the potential to uncover convergent or divergent synaptic phenotypes that may be specifically associated with autism versus schizophrenia risk. The imaging and overall breadth of the methods are convincing. The purely in vitro nature of the study slightly limits the generalisability of the findings, though these limitations are acknowledged and discussed in the manuscript.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Kashiwagi et al. undertook a population analysis of dendritic spine nanostructure applied to the objective grouping of 8 mouse models of neuropsychiatric disorders. They report that spine morphology in cultured hippocampal neurons shows a higher similarity among schizophrenia mouse models (compared with autism spectrum disorder (ASD) mouse models) and identify an effect of Ecrg4 (encoding small secretory peptides) on spine dynamics and shape in these models.

      Strengths:

      The study developed a method for objectively comparing spine properties in primary hippocampal neuron cultures from 8 mouse models of psychiatric disorders at the population level using high-resolution structured illumination microscopy (SIM) imaging. This novel technique identified two distinct groups of mouse models according to the population-level spine properties: those with ASD-related gene mutations and those with schizophrenia-related gene mutations. Functional studies, including gene knockdown and overexpression experiments, identified an effect of Ecrg4 on the spine phenotype of the schizophrenia model mice.

      Weaknesses:

      The main weakness is that the study is wholly in vitro, using cultured hippocampal neurons. The authors present this as an advantage, however, arguing that spine morphology as measured in a reduced culture system can demonstrate direct effects of gene mutations on neuronal phenotypes in the absence of indirect influences from nonneuronal cells or specific environments.

    3. Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for timelapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. My major concerns from the initial manuscript, especially regarding cherry picking and circularity have been addressed with revised analytical approaches. I have some remaining minor comments.

      (1) The comparison between two wild-type samples versus wild-type-mutant samples is helpful - I think this could be added to the manuscript.

      As suggested, we added the figure comparing two wild-type samples against wild-type mutant samples as Supplementary Figure 2. 

      (2) For results of time-lapse imaging - please spell out in the results section the direction of change (lines 270 - 277).

      As suggested, we added the direction of change (an increase in the turnover rate) to the text (page 12, lines 270-271).

      (3) Using linear mixed effect models for statistical analysis is a significant improvement. While a sample size (n) of mice = 3 is not ideal, I think given the multiple different mouse lines used and intensity of analysis, this is probably the best that can be done, although further validation in larger samples eventually is to be hoped for.

      We appreciate the reviewer for recognizing the effort required to collect data across multiple mouse lines.

      (4) The revised text is much improved, but I still think the authors should be upfront somewhere in the text that the schizophrenia-associated genes can only confer biased risk for schizophrenia (and that the clinical phenotype can also include autism). As I said before, I think this is the best we can do and I agree with their choices, but it is important not to overstate the link. The differences they see make it clear that these are still relevant distinctions.

      As suggested by the reviewer, we further modified the discussion related to the comparison between ASD- and schizophrenia-associated mouse models (pages 23-24, lines 508-522).

      “The nanoscale features of dendritic spines in mouse models of Nlgn3<sup>R451C/(y or R451C)</sup>, Syngap1<sup>+/−</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, which we classified as being related to ASD, are highly heterogeneous. This heterogeneity may reflect the broad clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. Accordingly, these four mouse models may represent distinct subgroups characterized by different degrees or forms of hippocampal dysfunction. Notably, among the ASD-related models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those found in the 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> mouse models. Although we classified 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> as schizophrenia-related models, both 22q11.2 deletion syndrome and Setd1a haploinsufficiency in humans are also associated with ASD, suggesting substantial overlap in the genetic risk factors underlying ASD and schizophrenia. Further systematic analyses linking rare genetic variants to synaptic phenotypes in mouse models may provide important insights into the mechanisms underlying both shared and disorder-specific synaptic alterations in neurodevelopmental and psychiatric disorders.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I would suggest that it might be preferable to use the word 'neuropsychiatric' rather than 'mental' in the title.

      As suggested, we modified the manuscript title.

      (2) I think it would be clearer to say that DEGs are listed if present 'in three or more models' rather than >2 (I appreciate the latter is mathematically clear, but can easily be read as 2 or more if reading fast). This is changed in the figure legend, but I suggest it is also changed in the main text (line 352-3)

      As suggested, we changed the main text to incorporate "in three or more models" (page 16, line 352).

      (3) Please add to Methods (line 557) that 'control cultures were prepared from littermate embryos....'

      As suggested, we added the phrase "control cultures were prepared from littermate embryos" (page 26, line 559).

      (4) Sorry to add something, but please could the authors add a definition of how they calculate spine turnover (and add units to the y axis of Figure 5A-C)?

      As suggested, we modified the y-axis of Figure 5A-C (% as unit) and added the method of calculating spine turnover rate in the text (page 36, lines 808-811).

    1. eLife Assessment

      This useful study combines experiments and mathematical modeling to show that antibiotic protection provided by resistant cells can extend across both surface-associated and freely growing bacterial populations. Notably, they show that treatment efficacy depends on population composition and density. The evidence supporting the main conclusions is incomplete, primarily because the biofilm context is not adequately characterized and demonstrated, raising the concern that it might represent only an aggregate of cells on the surface (rather than a biofilm) under the studied experimental conditions.

    2. Reviewer #1 (Public review):

      Summary:

      This important study examines how antibiotic-resistant bacterial cells can protect neighboring sensitive cells in mixed populations that occupy both surface-associated and freely growing states. Using experiments in Enterococcus faecalis together with a mathematical model, the authors test the hypothesis that protection would be stronger in biofilm-associated populations, but instead find that resistance-mediated protection extends broadly across both population types. The work provides evidence that antibiotic efficacy depends strongly on community composition, population density, and density-dependent detoxification dynamics.

      Strengths:

      A major strength of the study is the close integration of experimental measurements with a relatively simple quantitative model that captures many of the observed population dynamics. In particular, the work highlights how interactions between antibiotic detoxification, cellular growth, and saturation at carrying capacity can generate nonintuitive behavior, including the reported population inversion effect. The agreement between the well-mixed model and the experimental observations is convincing, and the spatial analyses suggest that cells within the biofilm are sufficiently intermixed that large-scale spatial segregation is unlikely to dominate the observed behavior.

      Weaknesses:

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The study also raises interesting questions regarding the role of spatial structure and exchange between planktonic and biofilm-associated populations. It would be informative to explore whether biofilm-specific protection becomes more pronounced at lower antibiotic concentrations, where local detoxification may compete more directly with antibiotic penetration into the biofilm, and in this context, the dynamics of exchange between biofilm and planktonic populations would be interesting to understand. Overall, the evidence supporting the central conclusions is convincing, and the study will likely be of broad interest to researchers studying microbial communities, antibiotic resistance, and collective population dynamics.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Martins et al. examined the cooperative response of E. faecalis cells to beta-lactams, in both planktonic culture and in biofilm. They found that the competition outcome between the susceptible and resistant strains is frequency dependent; they have also quantified how the competition curves change with inoculation OD and antibiotic concentration. To the authors' surprise, the competition dynamics are not that different in biofilm and in planktonic culture, which the author attributed to the unstructured nature of the thus-grown E. faecalis biofilms, quantified through correlation analysis. Using a well-mixed model capturing growth, death, and drug degradation by the resistant cells, the authors were able to quantitatively capture the experimental observation.

      Strengths:

      Overall, the data presented are solid. Although there is not much surprise after the understanding that the E. faecalis biofilm is unstructured, the manuscript still provides a useful "null case", so to speak, for researchers in the field when considering antibiotics in the context of biofilm. The theoretical model presented and the procedure of fitting the experimental data are useful to the research community.

      Weaknesses:

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

    4. Reviewer #3 (Public review):

      Summary:

      The authors studied social aspects of antibiotic resistance by co-cultivating antibiotic-resistant and sensitive Enterococcus faecalis (an important pathogen) as biofilms to assess the extent to which sensitive cells can take advantage of the protection provided by resistant cells against both a beta-lactam antibiotic and in the presence of a B-lacatamase inhibitor. By quantifying the proportion of each cell type using fluorescence microscopy, they conclude that protection is provided equally in the biofilm and planktonically, and that the biofilm is completely unstructured with regard to the locations of the two cell types. A mathematical model is then used to show that no spatial information is needed to recapitulate the results and that the protective effect can be described completely by the growth rates of the two cell types and the affinity of the β-lactamase to the antibiotic and inhibitor. The strength of evidence is difficult to assess due to unclear descriptions of some methods, and the significance of the findings is limited by the experimental setup, where antibiotics were added very close to the time of inoculation.

      Strengths:

      The co-cultivation of antibiotic-resistant and sensitive bacteria allows for exploration of the social aspects of antibiotic resistance. Fluorescently-tagged strains allow for unambiguous tracking of the two cell types. The simultaneous analysis of biofilm and planktonic cells enables insight into whether these different growth modalities are influenced by social aspects of antibiotic resistance. In analyzing the structure of the biofilm, the use of a null model with randomized cell positions allows for an accurate determination of whether the observed data are due to some effect; however, as noted below, there is a caveat to this analysis. The broad observation that biofilm and planktonic populations are linked is generally supported by the data; however, this result is closely tied to the experimental setup used. The development of a mathematical model that can recapitulate results from a second set of data with values obtained from fitting a different set of data shows robustness of the model for using it to explain the results.

      Weaknesses:

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. The described 'population inversion' effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed. Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account. The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case. The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

    5. Author response:

      We would like to thank the editors for their interest in our work and the three referees for their time and careful reading of the manuscript. The reviewers have provided a series of helpful suggestions that we discuss in this provisional reply and will seek to address in the revised version of the manuscript.

      The main concern raised is that the bacterial community we refer to as a biofilm may instead correspond to a cell aggregate. Following the passing of Prof. Kevin Wood, in whose lab the experimental work was carried out, our ability to perform additional experiments is limited. Nevertheless, we plan to wash and fluorescently stain the extracellular matrix before imaging to measure the extent to which the observed bacterial community is an attached biofilm. In the meantime, we would like to highlight the work of Wen Yu et al. [1], in which E. faecalis biofilms were grown in 96-well plates under antibiotic stress. In particular, one of the strains of E. faecalis used in this article was OG1RF, the same strain used in our study. Crystal violet staining was used to quantify biofilm biomass, providing evidence for biofilm formation under those conditions. While we recognize that the experimental setup differs from ours and that the OG1RF sample used did not contain fluorescent and resistance plasmids, these results nevertheless support the expectation that OG1RF will readily form biofilms.

      Reviewer #1 (Public review):

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The timescale of drug degradation is an important system metric. We appreciate the referee’s suggestion to quantify antibiotic concentration over time. We plan to perform experiments in which samples are collected from the culture at fixed time intervals. After removing the bacteria from the samples via centrifugation, serial dilutions of the supernatant will then be spotted on a lawn of sensitive cells to measure the antibiotic efficacy at each time point.

      Reviewer #2 (Public review):

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

      We agree with the referee that additional evidence would strengthen our study. As noted above, we will perform additional experiments to demonstrate the presence of an attached biofilm.

      Reviewer #3 (Public review):

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. [...] The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

      The experiment was designed to address how coupled planktonic and biofilm populations develop in the presence of antibiotics, which we will more explicitly discuss in the revised manuscript. We do agree that investigating how mature biofilms and their planktonic populations respond to antibiotic stress is an exciting direction for future studies. However, we believe that is beyond the scope of the our study on the development of coupled populations. We will be sure to explicitly identify this limitation in our revisions.

      The described ‘population inversion’ effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed.

      The reviewer is correct that this ‘population inversion’ is a frequency-dependent (perhaps also density-dependent) effect, and we should have situated it within that broader ecological framework. We will use this terminology in our revisions. We do want to acknowledge that the late Dr. Kevin Wood was fond of this phrasing to describe the reversal of the dominant strain, which is not necessarily true for frequency-dependent effects. Although we do not know for certain, we suspect this was a play on the ‘population inversion’ term used in quantum physics, used to describe a system in which its excited state (high energy) population unexpectedly outnumbers its ground state (low energy) population.

      The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case.

      If the reviewer is indeed discussing Figures 1a and 1b, these are not comparable as the starting fractions differ. On the other hand, if the reviewer was talking about Figures 2a and 2b (which is more clearly discussed by looking at Figures 2c and 2f), we agree that it appears that planktonic communities tend to have a slightly greater frequency of sensitive cells than the biofilms. We will be sure to highlight this observation and possible explanations in our revisions. However, given the uncertainty in these observations, we do not believe the differences are sufficient to alter our overall conclusion that final resistant fractions in biofilm and planktonic populations are quantitatively similar. Furthermore, the no drug treatment shows the same trend, which suggests it’s an effect of different growth dynamics of these two strains at high density rather than driven by the protective effects of resistance cells.

      Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account.

      The entirety of the Z stack was used to measure the final resistant fraction in the biofilm. On the other hand, we used only the densest slice of the Z stack to calculate the correlations. The correlations follow the same trend when calculated over less dense slices, but as density decreases, noise increases, so such plots did not bring more clarity to our conclusions and were not included in the manuscript. Additionally, only horizontal correlations (over a slice) were calculated because consecutive Z-stack slices were imaged with a 2.5 µm spacing. Given that the average cell diameter is approximately 1 µm, calculating vertical correlations may miss neighboring cells located between imaged slices, making such measurements unreliable. We will clarify the points raised by the reviewer in the results section and add more detail to the imaging methods section in the revised manuscript.

      References

      (1) Wen Yu, Kelsey M. Hallinen, and Kevin B. Wood. “Interplay between Antibiotic Efficacy and Drug-Induced Lysis Underlies Enhanced Biofilm Formation at Subinhibitory Drug Concentrations”. In: Antimicrobial Agents and Chemotherapy 62.1 (Dec. 2017), 10.1128/aac.01603–17. doi: 10.1128/aac.01603-17. url: https://journals.asm.org/doi/10.1128/aac.0160317 (visited on 01/11/2026).

    1. eLife Assessment

      This important study provides the first in vivo evidence that nonsense-mediated mRNA decay (NMD) in mature astrocytes regulates astrocyte function, synaptic plasticity, and anxiety-related behavior. Using a broad range of approaches, the authors show that conditional deletion of Upf2 alters astrocyte morphology and calcium signaling while impairing synaptic transmission and plasticity, providing solid support for the central conclusion that astrocytic NMD influences neural circuit function. Some key mechanistic claims remain incompletely supported, including whether phenotypes reflect astrocyte remodeling versus loss, the interpretation of synaptic engulfment data, the link between NMD targets and calcium signaling, and the extent to which calcium dysregulation explains the observed synaptic and behavioral effects.

    2. Reviewer #1 (Public review):

      Summary:

      Lituma and colleagues investigate the role of NMD in astrocytes, an underexplored question given that prior work on NMD in the brain has focused exclusively on neurons. Using a tamoxifen-inducible, astrocyte-specific Upf2 conditional knockout (cKO) mouse, they report that loss of astrocytic NMD causes: (1) reductions in astrocyte cell volume and surface area across hippocampus, visual cortex, and prefrontal cortex; (2) decreased excitatory synapse density, reduced dendritic spine density, and impaired synaptic engulfment; (3) deficits in basal synaptic transmission and LTP, with selective impairment of mGluR-LTD; (4) elevated spontaneous calcium transients in astrocytes; and (5) anxiety-like behavior in the elevated plus maze (EPM) and contextual fear conditioning paradigms. Transcriptomic analysis of FACS-isolated astrocytes identifies 277 differentially expressed genes, ~40% of which carry canonical NMD-inducing features, implicating pathways linked to calcium signaling, phagosome formation, and glial development. A rescue experiment using the CalEx calcium extrusion pump demonstrates partial restoration of synaptic strength and anxiety behavior when astrocytic calcium is normalized.

      The study addresses an important gap in our understanding of RNA regulation in glial cells, and the overall conceptual framework is well described. The experimental design is generally appropriate, and the multi-pronged approach lends the main claims a degree of validity.

      Strengths:

      (1) Novelty: This is the first study to systematically examine NMD function in astrocytes in vivo. The identification of astrocytic NMD targets via RNA-seq combined with an NMD-inducing feature classifier is a meaningful methodological contribution.

      (2) Multi-method approach: The authors combine morphological analysis (Imaris 3D reconstruction), synaptic markers (PSD-95, LAMP2 engulfment assay), spine density measurements, acute slice electrophysiology, two-photon calcium imaging, behavioral testing, and transcriptomics. The convergence across these methods strengthens confidence in the claims.

      Weaknesses:

      (1) While the transcriptomic analysis is a valuable addition, the connection between specific NMD targets and the observed calcium phenotype remains largely correlational. The authors identify Gabbr2 and Adora1 as upregulated candidates with canonical NMD features and speculate that their elevated expression drives aberrant calcium signaling. However, no validation (e.g., qRT-PCR or protein-level confirmation) of these candidates is presented. The mechanistic pathway between NMD disruption and elevated calcium is thus inferred from pathway analysis rather than demonstrated. This is a significant gap between the transcriptomic and physiological arms of the study, and the authors should be more explicit about this limitation or, ideally, provide at least one validated target.

      (2) The reduction in astrocyte surface area in cKO mice is interpreted as contributing to reduced synapse contact and engulfment capacity. This is a reasonable hypothesis, but the study does not directly demonstrate that reduced astrocyte territory correlates with reduced synaptic coverage at the level of individual cells or brain regions. The temporal sequence of these events is unknown. Do morphological deficits precede synaptic changes? Clarification and qualification of this causal chain in the Discussion would strengthen the manuscript.

      (3) LFS-induced LTD is unaffected, while mGluR-LTD is reduced. This is intriguing and potentially informative about astrocyte contributions to distinct LTD mechanisms, but the difference receives limited discussion. Given the relevance of mGluR signaling to calcium dynamics and the identified pathway enrichments (GPCR signaling), this specificity deserves more attention.

      (4) The CTRL + CalEx condition is included in the EPM experiment but not in the electrophysiology or calcium imaging experiments, making it difficult to fully assess whether CalEx itself has off-target effects on synaptic transmission or anxiety in wild-type animals. The CTRL + CalEx EPM data (Figure 7F) appears to show a modest reduction in open arm time relative to CTRL, which, if robust, would suggest that excessive calcium reduction in astrocytes is also anxiogenic. This finding would be physiologically relevant and deserves comment.

    3. Reviewer #2 (Public review):

      Astrocytes are highly responsive to their environment and play a range of critical roles in brain function. Lituma et al. theorize that one mediator of that responsiveness is the regulation of RNA stability. They therefore undertake an assessment of astrocytes missing Upf2, a protein required for mRNA degradation via nonsense-mediated decay. This is an interesting study, approaching astrocyte biology from a novel angle. The authors take on an ambitious set of experiments, spanning morphological assessment, synaptic engulfment, electrophysiology, behavior, and calcium imaging.

      The authors show convincing data that knocking out Upf2 in astrocytes impairs synaptic plasticity, affects behavior, and changes the complement of astrocytic mRNA. These results, in and of themselves, are intriguing and suggest that NMD is an important biological process in astrocytes, warranting further study.

      My primary concern is whether the authors may be largely studying dying cells. The idea that NMD disruption has a dramatic effect on astrocyte morphology is an intriguing idea, but it is not fully established here. The nuclei in the example cKO morphology images appear small and/or fragmented. This raises concerns that the authors did not ensure that they had the full 3D morphology of the astrocyte in the section, and the cell is in part cut off, which would compromise any data on the morphology. The authors state that the tissue was sectioned at 70 um. The diameter of an astrocyte in the adult mouse brain is typically between 50 and 70 um. Unless astrocytes are perfectly positioned in the center of the slice, at this thickness, the majority of astrocytes will almost certainly be partially cut off. More detail on how cells were chosen and what quality control metrics were implemented would alleviate concerns here. An alternative possible explanation for these small/fragmented nuclei is that cKO astrocytes may be unhealthy to the point that they are actively dying. Using the transgenic ZsGreen label, the authors state that they observe a size change (Figure S4); this is not readily apparent and is not quantified in any way. It does appear from these images that there may be a loss of some astrocytes; cell death, which would also be an interesting finding, is a fundamentally different process than morphologic restructuring in living cells. The authors do attempt to count astrocytes (Figure S6B), but do so with GFAP. This is a fundamentally flawed approach. Because GFAP is not readily detectable in most healthy astrocytes in most gray matter regions, GFAP should not be used to quantify astrocyte numbers; this experiment should be repeated with a better marker, such as Aldh1l1, Sox9, etc.

      Synaptic engulfment: This is an extraordinarily high degree of engulfment in the control animals compared to many published studies, leading to concern as to the technical approach. Indeed, the overall low level of PSD-95 signal in control conditions in adult mice is concerning as to the technical accuracy of the approach. It is unclear exactly how the investigators labeled the astrocytes; presumably via the ZsGreen label, but it is never stated, and the only images shown are the highly processed Imaris renderings. The small astrocytic processes, or leaflets, that make up the vast majority of the astrocytic arbor are on the order of 100nm in diameter. The processes shown in Figure 2B are, according to the scale bar, at least 20x that size. It is difficult to have much faith in these results as currently presented.

      The signal-to-noise ratio of the GCaMP experiments is worryingly low, likely responsible for the abnormally low dF/F in all conditions and the lack of significant change between control and CalEx, when control astrocytes should show a much higher GCaMP signal than any CalEx-expressing astrocyte. That said, the higher Ca++ in Upf2 KO astrocytes is intriguing. Given the roles of elevated calcium in cell death, this may reflect cells that are unhealthy to the point that they are starting to die.

      The authors conduct a FACS-based analysis of astrocytic mRNA from control vs Upf2-KO, with intriguing results. An important caveat, though, is that a large amount of astrocytic mRNA is in the processes. If mRNA stability is being actively and rapidly regulated, it seems likely that the mRNA in the processes would be the most relevant population of regulated mRNA. FACS-based approaches to astrocyte purification will, as robustly shown elsewhere, strip off those processes. Particularly given that the authors have shown that the processes may be the most actively changing astrocytic compartment with Upf2 KO, this is a strange choice of technique vs. something like Ribotag that would preserve the mRNA in processes. At least, there should be some discussion regarding using FACS for this analysis and the consequences for profiling mRNA in astrocytic processes.

      Minor points:

      (1) The use of the Aldh1l1-CreER mouse is a strong choice and has been shown to be highly astrocyte-specific. Combining that transgenic mouse with viruses driven by different forms of the GFAP promoter is quite bizarre in several ways. First, GFAP-dependent AAVs have been shown repeatedly to have significant neuronal leak. Second, these mice are, in all cases, receiving two different viruses, driven by different forms of the GFAP promoter, and the non-Cre virus is not Cre-dependent (vs. a much more standard approach of using a Cre-dependent second virus to ensure that all analyzed cells received both viruses). The authors mention that "this experimental design ensures that phenotypes are not caused by an acute effect of tamoxifen." It is certainly true that tamoxifen is not a biologically neutral molecule. However, the mice still receive tamoxifen, both in these morphology virus experiments and in almost all other experiments. This experimental approach is not inherently bad, nor does it necessarily invalidate the data (although the near-certain neuronal contamination due to the GFAP promoter-driven viruses is a concern). It is, however, convoluted in ways that appear unnecessary. If there is a strong rationale for this approach beyond the tepid explanation already present, it should be explicitly mentioned.

      (2) The characterization of the knockout is incomplete. While the authors should be applauded for their attempts to phenotype the cells in which they observe Cre-mediated recombination, there are issues with their technical approach. Most importantly, and an issue that affects other analyses in the paper as well: the vast majority of astrocytes in the healthy cortex do not express GFAP. Therefore, using GFAP to claim high astrocyte specificity and efficiency is a fundamentally flawed approach. Second, MBP is a myelin marker, not a cytoplasmic marker, and would not successfully colocalize with a cytoplasmic marker like ZsGreen even if recombination in oligodendrocytes did occur. Third, recombination at one set of LoxP sites is not a reliable indicator of recombination at other sites. Recombination efficiency is highly dependent on the spacing between the LoxP sites and cannot be reliably extrapolated to other floxed genes without validation. Finally, the most likely culprit for off-target recombination with Aldh1l1-CreERT2 (or other astrocyte-selective Cres, and certainly the GFAP-based viral promoters) is neurons, which the investigators did not test for. Neuronal Aldh1l1-CreERT2 leak is most likely to occur in the hippocampus. With the images shown in Fig S3, it is unclear whether it is possible to convincingly colocalize Upf2 staining with a cytosolic marker of all astrocytes, such as Aldh1l1 or S100b, but such data would be more appropriate. An alternative approach to validation would be in situ hybridization.

      (3) Supplementary Table 2 should include gene IDs, not just Ensemble IDs.

      (4) It is not fully clear what the investigators are denoting as a spine in Figure 2E; the two images do not appear to have the large degree of difference that the quantification suggests. The oversaturation of the signal complicates assessment.

      (4) A more detailed discussion of the rationale behind the timeline would be helpful. What is the half-life of Upf2, and how rapidly do NMD genes build up upon Upf2 disruption? In particular, in the case of virus experiments, the timeline is quite fast: ~2.5 weeks from injection to analysis. ssAAV expression takes over a week to reach appreciable levels.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigate mRNA targets of the nonsense-mediated decay (NMD) pathway in astrocytes and link the dysfunction of NMD in astrocytes to aberrant synaptic transmission that has downstream effects on behavior. Specifically, they find a link between the aberrant synaptic transmission with elevated spontaneous calcium signaling in astrocytes, and functionally they demonstrate that manipulating astrocyte calcium signaling with CalEx modulates astrocyte calcium signaling towards wildtype levels and improves anxiety behavior. They investigate the astrocyte calcium signaling changes in Upf2 conditional knockout mice in several brain regions that have been linked to anxiety behavior, including the hippocampus and prefrontal cortex. They also observe aberrant astrocyte calcium signaling in the visual cortex, demonstrating that dysfunction of the NMD pathway in astrocytes has widespread effects on synaptic transmission in various brain regions. This work identifies, through RNA-Sequencing, potential mRNA targets of NMD in astrocytes, and shows that pathway enrichment of these targets highlights calcium signaling. Altogether, this work highlights the importance of the basic cellular process of NMD in astrocytes, which are known to have extensive local translation of proteins in their perisynaptic processes. NMD may be particularly important in astrocytes due to their intimate association of processes with neuronal synapses, and the authors suggest that alterations to NMD function in astrocytes may be an important avenue for future investigation in neurodevelopmental disorders.

      Strengths:

      Altogether, this work is a critical foundation for future research into astrocyte contributions to neurodevelopmental disorders. The authors do a thorough characterization of astrocyte conditional Upf2 knockout mice in several brain regions. They present a complete story that connects molecular events (NMD pathway regulation of mRNA degradation) to astrocyte regulation of circuit activity to organismal behavior. The electrophysiological analysis is thorough, and the manipulation of calcium activity ties astrocyte calcium activity to anxiety behavior. The RNA-sequencing dataset is useful to the scientific community and provides a resource of candidate molecules that might be dysregulated in neurodevelopmental disorders.

      Weaknesses:

      The study suffers from some overstated claims and a lack of statistical rigor in some experiments, as detailed below.

      (1) The title states that "Astrocytic Nonsense-mediated mRNA decay regulates calcium signaling to support synapse function and restrain anxiety". The term "restrain anxiety" implies that the NMD pathway has a direct effect on a molecular switch to control anxiety. Anxiety behavior is a complicated process, controlled by many biological phenomena and synaptic transmission in the circuit as a whole, and is not directly linked to a specific NMD mRNA target. This title is overstating the findings of the study.

      (2) In general, the first figures (1-2) suffer from low power (N = 3) and statistical rigor. The statistics are inflated by analyzing individual fields of view and per-cell data rather than performing the statistics on the average of biological replicates. It is preferable to show the biological replicate data so that readers can observe the natural biological variability between replicates.

      (3) The claim that astrocytes have decreased engulfment of synapses in the Upf2 conditional knockout mice is not strongly substantiated by the data. The resolution of confocal microscopy and the static nature of histological images make it difficult to measure synaptic engulfment as an active process. Additionally, the metric of quantifying the % occupancy of PSD95 puncta within the total astrocyte volume may be skewed due to overall differences in cell size (shown in Figure 1). There is not much discussion of how a decrease in astrocyte engulfment of synapses may lead to decreased synapse number. To the contrary, one might expect decreased engulfment to result in increased synapse density.

      (4) The authors use Gfap as a marker to count astrocyte cell number and assess if there are changes in cell number between genotypes (Figure S6). However, Gfap does not label all astrocytes in the cortex and, in fact, is rather an aberrantly expressed marker in conditions of inflammation, as opposed to the hippocampus, where Gfap is basally expressed in all astrocytes. In the cortex, there seems to be a trend for reduced Gfap in the conditional knockout mice, which may suggest differences in astrocyte molecular signatures rather than cell numbers. Another astrocyte marker, like Aldh1L1, will be more accurate to assess this question histologically.

      (5) The authors state that "Preventing abnormally high basal calcium activity in NMD-deficient astrocytes restores normal excitatory synapse function...". However, this claim is not substantiated by the data. CalEx manipulation certainly shifts the input-output curve but does not restore to wildtype baseline levels (Figure 6E). Additionally, synapse number does not appear to be restored to wildtype levels (Figure 6D - although the p-value for this comparison is now shown). The investigators do observe improvements in anxiety phenotypes, suggesting there is some modulation of circuit activity, but the claim that CalEx manipulation restores baseline synaptic transmission is not supported.

    5. Author response:

      We thank the reviewers for their careful reading of our manuscript and for providing positive, constructive feedback. In particular, we thank the reviewers highlighting the several strengths of our study.

      To address the reviewers’ major concerns, we will revise the presentation of our main findings (specifically data/animal vs data/ROI), provide more clarity in the Results, Methods, and Discussion sections, and modify the title to better reflect these nuances.

      Additionally, we will perform the following new experiments:

      (1) Astrocytic Marker Validation: To further confirm comparable astrocyte cell counts between the CTRL and Upf2-cKO conditions, we will perform immunostainings using Aldh1L1, Sox9, or S100b instead of GFAP.

      (2) NMD Candidate Validation: To validate top candidate NMD target transcripts, we will perform immunostainings or qRT-PCR for Gabbr2, Adora1, S100b, or Cldn9.

      (3) Sample Size Expansion: To strengthen the morphological and PSD-95 quantifications, we will increase the sample size (N) by incorporating additional animals.

      (4) Mechanistic Timeline & Phenotype Linkage: We value the reviewer’s comment regarding the timeline of morphological and Ca<sup>2+</sup> phenotypes. To gai insight into whether these phenotypes are independent or linked, we will perform 3D reconstructions in CalEx conditions to assess whether Ca<sup>2+</sup> restoration rescues astrocyte morphology in Upf2-cKO mice. This will allow us to determine if increased Ca<sup>2+</sup> activity is upstream of the morphological alterations. Taken together, we believe that incorporating these manuscript revisions will strengthen the clarity and conclusions of our work. We thank the reviewers for their time and careful evaluation of our study.

    1. eLife Assessment

      This is a valuable paper looking at nanoscale organization of the membrane associated periodic cytoskeleton in mouse sciatic nerve axons. Despite previous studies, the precise organisation of the structure remains unclear, especially in vivo, and this manuscript significantly adds to this knowledge with solid data. An unexpected observation is the presence of discrete nanoscale clusters, regularly distributed around sections of axons. However, the paper misses a description of these clusters along the longitudinal axis of the axons.

    2. Reviewer #1 (Public review):

      Summary:

      The article "Nanoscale organization of beta-II spectrin within segments of the membrane-associated periodic skeleton in mouse sciatic nerve axons" by Gazal et al. looks into the organization of the spectrin scaffold in mouse sciatic nerves using super-resolution microscopy. It is now well established that axons, across species, contain a membrane-associated periodic scaffold mainly composed of circumferential actin filaments and longitudinally arranged spectrin tetramers. While super-resolution imaging of neurons in cell culture is relatively easy, exploring the ultrastructure of myelinated axons in intact nerve fibers is a daunting task. Nevertheless, the authors have attempted this by fixing and preparing cross-sections of sciatic nerves. They have then tried to quantify the fluorescence intensity patterns of specific components, especially that of labeled beta-II spectrin and have analysed its distribution.

      One of the main findings is that spectrin is distributed along the axonal periphery and along the outer part of the myelin sheath. By labelling multiple cellular components and using intensity analysis, the authors show the sequence of structural organization of a few key components. They see that, unlike in the case of axons in culture, the axonal cross-sections within the sciatic nerve deviate significantly from a circular shape. They then use 3D-dSTORM to investigate the distribution of beta-II spectrin along the axonal circumference. They see that this distribution is very heterogeneous, both in the sizes of spectrin puncta and their arrangement along the periphery. The amount of spectrin scales linearly with axonal circumference.

      Strengths:

      Super-resolution imaging of axons of intact nerve fibers to investigate the organization of beta-II spectrin.

      Weaknesses:

      While most of the findings, like the spatial distribution of spectrin and related components, are reasonably well supported by data, I have concerns regarding the subsequent claims made in the article. The detection of axial periodicity based on the observation of a peak in the inter-tetramer spacing distribution is not very convincing, and a 3D representation (or a video of 3D reconstruction) would have been better. And so are the claims on characteristic spectrin spacing of 200 nm along the axonal circumference. A peak in the distribution does not imply a periodic arrangement.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting paper by the Unsain lab looking at the nanoscale organization of the membrane-associated periodic cytoskeleton in mouse sciatic nerve axons. The precise organization of the structure remains unclear, especially in vivo, and this manuscript significantly adds to our knowledge of this important structure. While some of the findings in the study are somewhat expected (though still valuable to see in an in vivo setting), an interesting observation is the presence of discrete nanoscale clusters that scale up with the size of the axon, which challenges previous assumptions.

      Strengths:

      Strong, convincing data; clever combination of imaging and analytical tools to make novel points; well written; excellent composition of figures.

      Weaknesses:

      (1) Figure 2A/3A: The large and small clusters of spectrin, as seen in cross sections, are unexpected and novel. The authors have done a clever job of combining imaging and analyses, but some things are still unclear. First, the authors should be consistent in their language when they talk about the spectrin clusters. Recommend precise language to define the small and large clusters when they first appear in the text, and then use the definitions consistently throughout the text. Second, based on the data shown, one does not get a clear idea of how the small and large clusters are organized along the longitudinal axis of the axon. In that context, are Figure 2B and C from imaging along the longitudinal axis? If not, it's unclear how the authors can conclude that the spectrin assemblies have a distance of ~170 nm along the linear axis. In general, a perceived limitation of this study is that while the authors have done a good job looking at cross sections, there is no information on the longitudinal distribution of spectrin in these axons. Looking at both cross- and longitudinal sections would also clarify details about the large spectrin clusters. For instance, are they small sausage-like structures, or long rods of spectrin running along the length of the axon? One assumes that all the analyses in Figures 3 and 4 are from the small clusters. Can the authors do a similar analyses of the large clusters? Finally, a schematic model showing both cross- and longitudinal- sections would make things clearer, but the authors would need to show the longitudinal data for that.

      (2) It is interesting to think that the larger spectrin accumulations may be similar to the condensate-like structures seen by Boyer et al., as the authors mention in the discussion. In that context, it is possible that these focal accumulations are local reservoirs of spectrin that are also seen in mature axons (indeed, these accumulations were also seen in mature axons in the Boyer et al. paper, and they also speculated that these accumulations may be local reservoirs). Can the authors check if actin/adducin is also present in these larger spectrin accumulations?

      (3) While talking about the nanoscale clusters, it is important to specify that the authors are talking about circumferential clusters. Though the writing is excellent, one still does not get the precise definition of "clusters" from just reading the abstract, and it would be good if the authors could work on that more (I recognize that this is not easy to do).

    4. Reviewer #3 (Public review):

      Summary:

      In the presented work, the authors investigate spectral staining in axons of the sciatic nerve, where the MPS has been detected before using STED microscopy. They employ 3D-dSTORM in tissue sections and analyze the data, measuring localization of clusters on the axon perimeter and the relative distribution of those. From these data the conclude that large gaps in spectrum localizations exist and that clusters around the axon exist that are spaced at 200nm.

      Major Comments:

      (1) The presented data are at times overinterpreted, and the discussion lacks a critical view of the data. For example, the statement "...Unlike previous suggestions from qualitative evidence in cultured neurons (REfs), βII‑spectrin distribution in MPS segments of peripheral nerves is discontinuous, with extensive stretches of the perimeter lacking βII‑spectrin." is quite strong, given it is based on immunofluorescence staining and dSTORM microscopy in tissue. Absence of evidence of staining is not evidence of absence.

      (2) The authors claim in the abstract that "The number of these clusters scales linearly with the axonal perimeter, maintaining a constant membrane occupancy of ~20% across varying axon diameters." Again, this is from a cut through an axon, while measuring the density of clusters on the perimeter. If they claim area occupancy, an area should be imaged, and the dots (clusters) should be measured in surface coverage in a 2D projection of the axonal surface.

      (3) In general, this reviewer suggests being a bit more moderate in statements such as: "These findings challenge simplified models of the MPS based on cultured systems and demonstrate that the MPS in peripheral nerves is composed of discrete structural units." These statements are bold from the relatively few measurements in a single method and a single viewpoint. Especially when considering that techniques such as dSTORM depend extremely highly on labeling density, and apparent clustering of localization is highly prone to misinterpretation. If the authors desire to make such statements, working with endogenously labeled protein would be warranted. The authors should at least hedge such statements.

      (4) If the authors want to make statements about general organization, why do they not compare adjacent cuts through the axon? If there are continuous spectrin filaments, the clusters should appear at the same site across repeated cuts through the axon.

      Besides this, this reviewer welcomes the effort that has been made to establish dSTORM in tissue sections and to investigate the MPS in native tissue.

    5. Author response:

      We sincerely thank the editors and reviewers for their overall positive assessment and constructive feedback on our manuscript detailing the nanoscale organisation of βII-spectrin of the membrane-associated periodic skeleton (MPS) in mouse sciatic nerve axons. Their perspective and comments will help refining the manuscript.

      A common comment by the reviewers relates to the description of the characteristic longitudinal periodicity of the MPS. We value these comments, which we believe are motivated by the fact that the longitudinal periodicity of the MPS is undoubtedly the most studied and prominent feature of the MPS in cultured neurons. However, the main goal of the present project was to describe how βII-spectrin is organised in the transverse axis of individual segments of the MPS in nerve tissue. This is why we utilised cross-sections of the sciatic nerve, hence achieving the best resolution possible in that plane, at the expense of the resolution in the axial axis. Furthermore, this study clearly shows that the transverse morphology of axons, and thus of the MPS, of neurons in the tissue is highly irregular, in comparison to cultured neurons. This imposes an extra challenge to observe correlated longitudinal structures when the observation length is limited, as in our studies. Nonetheless, to improve this aspect of the manuscript, we will revise our data and previous evidence, clarify the methodological trade-offs made, and make our interpretations more accurate.

      Additionally, we will clarify several imaging- and definition-related inquiries, including tests for insufficient staining, the interpretation of βIII-tubulin staining, the assessment of axon–glia boundaries, and consistency in the use of terms like ‘clusters’ and ‘periodicity’, among others.

      We believe these and other revisions will substantially strengthen the manuscript and comprehensively address the reviewers' feedback.

    1. eLife Assessment

      This valuable study shows the impact of the metabolic state of bacteria on phage infection. The experimental results, based on various phages infecting E. coli, are convincing and consistent with a two-step adsorption mathematical model. This study should be of interest to the communities working on cell metabolism and on host-pathogen interactions.

    2. Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. The mechanistic interpretation offered by the model highlights an interesting trade-off between adsorbing to a metabolically active host and discriminating between active and inactive hosts that, somehow, a phage has to optimize. It would be interesting, in the future, to investigate how different phages optimize this task.

      Comments on revised version.

      I am happy with how the authors have addressed all the comments. The manuscript is much clearer and more readable and the previous overstated claims have been removed/clarified.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phage that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well designed and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication and the paper shows that host physiology has to be taken into account at all steps of phage infection.

    4. Reviewer #3 (Public review):

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, 𝜙80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      As I have stated in the first submission of this manuscript, the data presented by the authors clearly supported the principal conclusion of the study. The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully use a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Comments on revised version.

      In this revised version, the authors have successfully resolved all of my comments. I appreciate that the main text has been majorly revamped, which greatly helps the readers follow the motivation behind the experiment and analyses, and interpret the data. I agree that the revised terminology choice "commitment to infection", instead of the previous interchangeably used "adsorption"/"entry", is much more logical, considering the experimental data. I also commend the authors for writing the modeling part in a very clear, pedagogical, and instructive manner. Overall, I believe that this manuscript will be valuable to those who are interested in phage-bacterial interactions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource-limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. While the model is sensible from a mechanistic perspective, the experimental evidence that supports how each model's rate is affected by the cell metabolic state is weak, as only ratios of these rates can be inferred from the data.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phages that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well-designed, and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication, and the paper shows that host physiology has to be taken into account at all steps of phage infection.

      Weaknesses:

      There are some concerns about the experimental setup and which conclusions can be drawn from it:

      Before phage infection, bacterial cultures are grown to exponential growth, washed, and then resuspended with glucose or arsenate-azide for 10min. It is however, questionable that 10 minutes is enough to simulate high and low metabolic states realistically. 10 minutes seems to be quite short to go from exponential growth to a low metabolic state, given the transcriptional memory of previous environments. It seems more likely that the population will be quite heterogeneous, with cells in various states of transition towards low metabolic states.

      While we agree with the reviewer that during metabolic transitions there may be a period in which the population is heterogeneous, with cells in different stages of transition toward a low metabolic state, the 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyper diffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027). We have also corrected the DOI for this reference in the manuscript. Furthermore, the ATP pool of log-phase E. coli turns over several times per second (Holms et al., Arch. Mikrobiol. 1972, http://dx.doi.org/10.1007/BF00425016). We therefore assumed the bacteria were energy depleted after 10 minutes. We have clarified this point in the revised manuscript.

      Given that arsenate and azide inhibit cellular metabolism, i.e., have antimicrobial effects, cells might not just downregulate metabolism but also activate the stress response, and this causes some of the observed effects on phage adsorption. Therefore, the 'low metabolic state' of the cells in this paper could mean that cells are starved or that they are stressed or both.

      The reviewer is correct. We don’t exclude indirect effects. However, as nutrients were removed from the bacteria by washing and energy metabolism was inhibited by the addition of arsenate and azide, we assumed a stress response requiring biosynthesis would be unlikely to occur.

      The abundance of receptors could change between the high and low metabolic media conditions and contribute to the observed differences in adsorption, while the authors seem to assume in their model that the initial adsorption rate always remains the same.

      We do not think that the observed differences in adsorption are explained by a change in receptor abundance. In a previous study using the same experimental protocol as in the present work, phage λ was compared to the metabolically insensitive mutant λh (Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119). If the lower adsorption in the low-metabolic condition were caused by a reduced number of receptors, then λh should also have shown a lower adsorption rate under the same condition. Instead, no measurable effect on λh adsorption rate was observed. We therefore conclude that the effect is not explained by changes in receptor number on the timescale of the experiment. We have clarified this point in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, ɸ80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy-depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      Strengths:

      The data presented by the authors clearly supported the principal conclusion of the study ("Viral commitment to infection depends on host metabolism"). The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully used a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Weaknesses:

      (1) The authors isolated and measured the numbers of free phages in the medium after infection of bacteria under different treatments. These measurements were analyzed in two different ways: (1) simply as ratios (corrected/normalized using different controls), and (2) fitted using a simple mathematical model. I have concerns regarding both analyses.

      (1.1) For the first method, having different time points at which the sample of each phage is collected critically complicates data interpretation. As one incubates the phage-bacteria mixture for a longer time, more infection occurs, and the number of phages collected from the mixture decreases. Therefore, the different incubation time forfeits the goal of "a systematic and quantitative comparison across different phages [...]", just as the authors self-criticized. Conceivably, the authors could have used the shortest measurement time for all phages (i.e., 10 minutes, as for phage λ). Alternatively, the authors could have applied a systematic criterion such as half (or any other fraction) of the latent period of each phage, which would still "maximize the incubation period while ensuring that manipulations were completed before the first infection cycle concluded". In my view, the seemingly arbitrary measurement time for each phage renders the entire first analysis very challenging to interpret. It also goes against the author's proposition that the protocol was "standardized" or "consistent". It is not clear what the readers are supposed to take away from this first analysis, or rather, which evidence, finding, or conclusion the manuscript would lose if the authors only presented the modeling-based analysis.

      (1.2) The second method of analysis sought to remove the dependence of the measurements on time. I completely agree with this goal, and the findings extracted from this analysis significantly contributed to the merits of this manuscript. However, the authors achieved this goal using a single time point for each phage to calculate the infection rate (η). As shown in Figure S3, each of the phage depletion curves is anchored by only one data point (note that the P(t)/P(0) = 1 at t = 0 is assumed, not measured). This goes against the typical way this collision model is used in the literature, where a time series is measured and used to fit the model (e.g., DOI 10.1007/978-1-60327-164-6 18, or more recently, PMID 39700139). This practice in the current manuscript reduced the robustness of the inferred η values. This problem is exacerbated by assumptions used by the authors in formulating this model. For instance, the authors used a constant value for the bacterial concentration, B, because "bacterial growth and lysis were negligible" (lines 135-136). However, considering that the bacteria were cultured at 37oC in a very rich medium (first in YT broth, then in 2% glucose), the measurement times of 20, 30, and 55 minutes are most likely one or a few generations of bacterial growth and division.

      Related note: I suggest that one of the panels in Figure S3 should be moved to the main text, since it is critical to the second method of analysis.

      We would like to clarify that the manuscript does not present two separate methods, but rather one method presented in two steps: a first step with results that are directly tied to the experimental measurements and show whether the effect is present for each phage, followed by a second, analytical step that makes the results comparable across phages.

      The first step presents the ratios because they directly reflect the measurements performed in the experiment and allow the reader to see the effect of the metabolic state for each phage in contrast to its control. We agree that these ratios are time-dependent and therefore not suitable for quantitative comparison between phages. Their purpose is to illustrate the experimental outcome and to show that the effect is present (or absent) on a per-phage basis not to compare magnitudes across phages.

      We then follow this with the second step, allowing the reader to follow the logic of the analysis. The analytical step that follows does not represent a second method, but a continuation of the same analysis. Here, we remove the time-dependence specifically in order to make comparison of the effect across phages possible, by connecting our results to standard measures such as the adsorption rate η. Importantly, P(0) is measured for every phage in every experiment. The only modeling assumption used (a standard one in the field) is the exponential form for the decay in free phage number, which naturally yields P(t)/P(0) = 1 at t = 0.

      Regarding the reviewer’s concern that bacterial growth may not have been negligible over the relevant time window, we note that recent work on rich-to-minimal growth lags in E. coli reports substantial delays before growth resumes after nutrient downshift. One 2023 study (Wu et al., Nature Microbiology 2023, https://doi.org/10.1038/s41564-022-01310-w) considering wild-type E. coli shows in Fig. 2c a lag of up to about 2 hours after a shift from MOPS minimal medium with 0.2% glucose plus 18 amino acids to the same medium without amino acids. Another 2023 study (Zhu and Dai, Nature Communications 2023, https://doi.org/10.1038/s41467-023-36254-0) examining both rel+ and rel− strains reports a growth lag of about 49 minutes for rel+ and more than 5 hours for the relA deletion strain. While these conditions are not identical to ours, they support the general point that growth does not immediately resume after such shifts. We therefore think it is unlikely that, following transfer from YT, the cells underwent one or a few full generations during the time window of our adsorption measurements.

      On the related note: Following the comments of all reviewers on Figure S3, we have decided to remove it to avoid confusion.

      (2) The data were able to distinguish phages that successfully infected bacteria and those that remained free in the medium, and the authors appropriately interpreted the data as such throughout the Results section. However, in the Discussion (starting from the very first sentence, line 172), the authors used terms that include "adsorption" and "entry" more interchangeably (for example, see the three sentences in lines 310-313, for "viral entry efficiency is shaped by [...]", then "adsorption kinetics modeling"). I do not see how the authors' data could distinguish between adsorption (the phage particles attaching to the outside of the cell) and entry (the phage DNA being injected into the cell). Conceivably, any phage particles that irreversibly attach to a cell but do not yet inject their genome into the cell would still be removed from the medium and therefore not quantified. Another example: in lines 189-191, the authors interpreted that "[...] when the bacterium is in a low metabolic state, the phage does not bind irreversibly to the host", but how do the authors eliminate the case of no phage binding (i.e., the reversible step) to begin with?

      We agree with the reviewer that our use of the terms adsorption, entry, and infection should have been more careful. Our experiment can only identify the irreversible commitment of phage to a host cell. We have therefore revised the text to refer consistently to phage commitment.

      Similarly, in lines 283-293, how do the authors delineate whether energy depletion would increase the k_off term or decrease the k_inj term, because either would result in more free phages in the medium as observed in the data? I believe that the writing of the Discussion, as it stands now, is doing a disservice to the conclusions presented in the Results section.

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: if energy depletion leads to reduced commitment, this can arise either because k_off increases, because k_inj decreases, or because both change, as long as k_off/k_inj becomes larger in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also leads to the trade-off now discussed in the revised manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become significantly larger in the inactive case. Depending on whether this is achieved through changes in k_off or k_inj, the cost of discrimination appears either as slower commitment or as additional energy dissipation. We agree that the previous wording overstated the mechanistic interpretation, and we have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (3) The authors presented an argument that performing infection of all five phages in the same condition is an advantage, allowing for comparison across different phages. While this goal is a completely valid one, it is difficult to reconcile that with the fact that different phages require different optimal conditions for successful infection. For instance, phage T5 famously requires Ca2+ for successful infection into the host bacterium (and later successful replication); see PMID 13174489. However, all infections were performed in TMG, which lacks Ca2+. Perhaps the absence of T5 dependence on the host metabolism is because the infection condition used by the authors was not optimal for T5 to begin with? Similar arguments could be made for other phages.

      Our study alone cannot eliminate that possibility. However, we have cited multiple previous studies, for example references citing Braun et al., showing that T5 remains insensitive to the host metabolic state under different buffer conditions. We therefore believe it is unlikely that the lack of metabolic dependence we observe for T5 is simply due to suboptimal infection conditions.

      (4) Whereas the manuscript examined five coliphages, only phage T5 and phage λ were discussed extensively. I believe some discussion points for these two phages need clarification.

      We focused our discussion on the phages T5, λ and φ80 because these are the phages for which similar effects have been reported previously in the literature. This allowed us to connect our findings directly to existing work and to discuss mechanistic hypotheses in a meaningful comparative framework. For the remaining phages, to our knowledge no prior studies have examined their behavior under comparable metabolic conditions, and therefore a similarly detailed discussion would have been speculative. Nevertheless, all five phages are treated equally in the presentation of the experimental results and in the quantitative comparison of adsorption rates.

      (4.1) Phage T5: The data obtained by the authors show that the infection rate of phage T5 is not impacted by the metabolic state of the host cell. Considering that the authors used the terms "infection", "adsorption", and "entry" interchangeably to refer to the irreversible commitment of a phage to a host cell (see point 2), this discussion regarding phage T5 lacks one critical literature context: DNA entry of phage T5 is known to occur in two phases (first-step transfer and second-step transfer). Critically, the second step can only occur if phage proteins encoded by the phage DNA transferred in the first step are expressed (see PMID 10577483 and the cited papers therein). In that context, metabolic poisoning of the host bacteria should have impeded T5 infection. The authors should comment on this point.

      As the reviewer pointed out, our usage of the terms infection, adsorption, and entry should have been more careful. Our experiment can only identify irreversible commitment of phage to a host cell. For T5, we expect that this irreversible commitment already occurs upon first-step transfer of phage DNA. As a result, even if second-step transfer is impeded under metabolic poisoning, our method would not resolve that effect. We have added this clarification to the revised manuscript.

      (4.2) Phage λ: The experiment using phage λ in this current study shares many resemblances to that in Brown et al. 2022. That feature alone is not a problem, but at many places in the text, the writing is ambiguous as to whether it is discussing the results in Brown et al. 2022 or in the current manuscript. I am giving three examples below, but this is not exhaustive: (i) Lines 67-69, there is no Brown et al. 2022 reference immediately after "a mutant phage variant (λh) could bypass this dependency [...]" (not just in the previous sentence); (ii) Line 228 should clearly say "Our previous findings suggested that phage λ is capable of [...]", since it concerns Brown et al., 2022, not the current study; and (iii) Lines 245-246, there is no Brown et al., 2022 reference immediately after "we observed that a mutant variant [...] even energy-depleted host" (without a reference, it reads like the authors "observed" that finding in this current manuscript).

      The reviewer is right. In those places, the text was ambiguous as to whether it referred to the present study or to Brown et al. (2022). We have now inserted the reference at the relevant points and revised the wording where needed to make this distinction explicit.

      Also, regarding phage λ: The discussion between line 230 and line 249 is very interesting, but since it concerns the differences between λ PaPa and Ur-λ, the authors should consider mentioning and discussing a very relevant recent study, PMCID: PMC6312755.

      We agree that the study by Guan et al. is very relevant and interesting. However, our point in this part of the Discussion is only to clarify that we used λ PaPa and not the originally isolated λ strain. We have therefore limited the discussion here to that distinction.

      (5) Control experiments, or references to prior studies, are needed to support that the As/Az treatment at this concentration and duration (at least 10 minutes) is sufficient to deplete the metabolic state of the cell. For instance, this can be shown by impeded or null cell growth, arrested motility (using a standard swimming assay), or a fluorescent reporter for the energetic state of the cell.

      The 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyperdiffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027) where the effect was assessed by monitoring the rate of movement of the λ receptor on the bacterial surface. We have clarified this point in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As mentioned earlier, I found the paper interesting and addressed an important and significant knowledge gap.

      My biggest concern is about the interpretation of the experimental data in light of the two-step model. In particular, around line 286, it is stated "k_inj is more sensitive to metabolic state than k_off". Assuming k does not depend on metabolic state, which is a fair assumption, the equation for eta only depends on the ratio between k_inj and k_off and not on the individual parameters separately. Consequently, there is no way of saying which one of the two is more affected by metabolic state, unless the model already assumes that k_off is not influenced by metabolic state. The results could equally be explained by k_inj decreasing in metabolically depleted cells, or k_off increasing in such cells. If this is an assumption of the model, this should be clearly stated and not reported as a consequence of the data, as it is at the moment. Also, how does this mathematical model connect to the fitting function used in Figure 2b?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: discrimination requires that k_off/k_inj be larger for inactive hosts than for active hosts, such that commitment is specifically reduced in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This introduces a trade-off: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination requires this ratio to become significantly larger for inactive hosts. If this is achieved through changes in k_off, discrimination comes at the cost of slower commitment by allowing more time to leave; if it is achieved through changes in k_inj, it can preserve fast commitment to active hosts but requires additional energy dissipation in order to actively modulate commitment. We have therefore revised the text accordingly to frame the argument in terms of this trade-off, rather than attributing the effect specifically to k_inj. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      I have a related experimental criticism. The kinetic model presented assumes an exponential decay of free phage, which is a commonly used assumption in the phage literature. Given that the phage types used in this study lyse relatively slowly, it would be good to actually see adsorption curves, in which free phage is measured at different time points between inoculation and lysis. This data would not only provide useful evidence for the kinetic model, but it should also replace what is now in Figure S3, which consists of fitting one experimental point with one line. As it currently stands, Figure S3 is not useful actually misleading.

      We appreciate the reviewer’s point. We agree that adsorption curves, in which free phage is measured at different time points between inoculation and lysis, would provide a stronger basis for evaluating the kinetic model. However, we do not have the resources to perform these additional experiments within the scope of the present study. Following the comments of all reviewers on this point, we have therefore decided to remove Figure S3 to avoid confusion.

      Finally, it is not clear to me why the quantity "Ratio" has been chosen to be presented in Figure 1, rather than the ratio of estimated adsorption rates eta'/eta, which is much more intuitive for a phage study and contains the same information. I would recommend switching to this choice, unless there is a clear rationale for why the quantity "Ratio" is more useful/effective. Showing eta'/eta would also increase the readability of Figure 1, as it would move the y-axis to a logarithmic scale and better visualize values around 1.

      We used “Ratio” in Figure 1 to illustrate the experimental design, controls, and measured quantities directly, as it more transparently reflects the data collected. In the second part of the analysis, where we compare time-independent adsorption rate estimates, we have presented the corresponding values of η′/η as suggested.

      Minor comments:

      (1) Introduction

      Line 31: "... such as nutrient limitation, fluctuating temperatures, and variable energy availability" - if drawing a distinction between energy availability and nutrient limitation, please make explicit what this distinction is. Energy availability seems like a natural consequence of nutrient availability.

      While energy and nutrient availability are often linked in E. coli, they represent distinct physiological constraints. Nutrient limitation refers to the lack of essential biosynthetic precursors such as nitrogen, phosphorus, or amino acids. Energy availability, in contrast, reflects the cell’s ability to generate ATP and reducing equivalents through metabolic processes. For example, under anaerobic conditions, E. coli may have ample nutrients but limited energy production due to the lower efficiency of fermentation compared to aerobic respiration. Thus, energy limitation can occur independently of nutrient limitation.

      (2) Results

      (a) Whole Section: Please label equations.

      All equations have now been labelled in the revised manuscript.

      (b) Lines 105 to 114: As stated in Major Comments, I think the clarity of the paper would be improved by introducing the relative adsorption rate here and dropping the concept of Ratio entirely. However, if the authors wish to use Ratio, I would recommend the following:

      Lines 105 to 109 are confusing to read because of the number of connectives: "... ratio of free viruses from permissive AND resistant hosts respectively TO the free viruses in buffer under energy-depleted AND energy competent conditions". This would be clearer if each quantity were given an algebraic symbol, and RP, RR, and Ratio were defined through formal algebra, rather than mixed mathematical and sentence notation.

      This section has been rewritten for clarity. We now introduce explicit algebraic symbols and define the quantities formally, which removes the ambiguity present in the sentence-only description while retaining the intended meaning.

      The chemical names "arsenate" and "azide" should appear in the body of the text before they appear abbreviated in an equation. Please state at this point that these are both metabolic inhibitors, as it is not immediately clear what role they play or why you are using them.

      The text has been updated to introduce arsenate and azide by name before the abbreviations are used, and we now explicitly note that they act as metabolic inhibitors.

      On line 114, the authors helpfully provide an interpretation of Ratio = 1. It would be useful to provide at the same time interpretations of Ratio >1 and <1, perhaps 2 and 0.5 specifically?

      We have added brief explanations illustrating the interpretation of Ratio values greater than and less than 1, including examples of 2 and 0.5.

      I would consider giving this quantity a more interpretable name than Ratio. This quantity represents how much a bacteriophage preferentially adsorbs to metabolically active cells, so perhaps "Selectivity" or "Adsorption Bias"?

      We intentionally retained the generic term “Ratio”, as this quantity reflects an intermediate experimental measure used to describe the process rather than a newly defined metric. Its purpose is to bridge the experimental observations and the subsequent quantification of effects on the adsorption rate (η).

      (c) Lines 117 to 122: the authors sometimes refer to ratios explicitly, "average ratio of around 1.6" and other times say e.g., "a greater than 3 times increase in viral particles". Using more consistent language (saying "Ratio" every time) would be clearer.

      We have standardized the terminology in this section and now refer to all fold-changes consistently using “Ratio” to avoid ambiguity.

      (d) Figure 1

      Phages λ and T6 look like they have ratios less than 1 for resistant cells? If this is true / if the ratio is statistically significantly below 1, please comment.

      The ratios for λ and T6 are not statistically different from 1. The apparent deviation is within the standard error of the mean. To make this clearer, we have added the corresponding p-values to Table S2 in the Supplementary Information.

      Ratios near 1 are difficult to distinguish from 1, especially in panels A and D. Using a logarithmic scale on the y-axis would make the plots more readable.

      Because the values in these panels are not statistically different from 1, changing to a logarithmic scale would not alter the interpretation. We therefore retained the current axis scaling to reflect that there is no meaningful deviation from 1 in these cases.

      The data corresponding to individual experiments have no error bars. Given that the number of free virions was determined by plaque assay, which carries an intrinsic sampling error, this uncertainty should be reflected in the plots.

      We thank the reviewer for this important comment. Because plaque assays have compound sources of stochastic variation, assigning a per-measurement error bar would risk implying false precision. For this reason, we present the values from each biological replicate directly, and the uncertainty is represented in the statistical summary across replicates. Specifically, for each phage and condition we show the three independent experimental measurements and report the mean along with the standard error of the mean. This approach allows us to represent biological variability without implying a precision that cannot be accurately quantified at the level of single plaque counts.

      Similarly, the average value does show error bars, but it is not stated what these error bars correspond to: standard error in the mean, standard deviation of the sample, or combined uncertainty?

      The caption has been updated to state that the error bars represent the standard error of the mean.

      The resistant bacteria seemed to have ratios close to 1 in all cases. Is this because very few virions adsorbed under both energy conditions?

      Resistance is commonly associated with a lack of a surface receptor for the phage (or generally an entry pathway). We use the resistant bacteria as a control group for the effect of the conditions on adsorption. For resistant bacteria, the Ratio should be 1 since virions do not adsorb under both energy conditions. Any slight variations from 1 should come from sampling errors or small heterogeneity in the population.

      (e) Figure 2

      Please comment on what the error bars here represent. Error bars in Figure 2 A seem to permit negative (or at least zero) values of relative adsorption rate for phages m13 and T6, possibly implying an overestimate of the error? If it is the case that multiple values used to calculate the mean are far apart, possibly showing the values individually through a superimposed swarm plot would be clearer.

      This point is now addressed in the Supplementary Information, where we clarify how the error bars were calculated.

      (3) Discussion

      (a) Line 189: "high metabolic state" is imprecise. Say "energy-competent" to be consistent with earlier language.

      To maintain continuity with earlier terminology, we now include “energy-competent” in parentheses alongside “high metabolic state,” while retaining the original phrasing for readability.

      (b) Figure 3, population level

      Show adsorbed virions physically attached to bacteria, rather than removing them completely from the image, as currently, the implication is that at a high metabolic state, there are fewer virions total, not fewer virions remaining in solution because more are adsorbed. You could go as far as to add a third "after centrifuging" row, showing the adsorbed phages stuck in the pellet and the unadsorbed phages remaining in solution.

      Thank you for this suggestion. Figure 3 has been updated to depict adsorbed virions attached to bacterial cells, clarifying that the decrease represents adsorption rather than loss of total particles. This change improves the accuracy and interpretability of the schematic.

      (4) Methods and Materials

      (a) Figure 5

      The step "estimate cell numbers from OD" appears to follow incubating plates overnight. If the cells you are counting come from the pellet produced by centrifuging 3 steps prior, you could add a fork into the black line connecting the steps, with one branch corresponding to the supernatant and phages, and the other to the pellet and cells?

      Thank you for pointing this out. The order in the figure has been corrected: cell numbers are estimated from OD before overnight incubation. This resolves the confusion without the need for branching in the workflow diagram.

      (a) Line 332

      You allow as much time as possible for adsorption without the possibility of lysis. Did you determine the lysis times / latent periods of these phages through one-step-growth-curves, or use published results, in which case please cite? Having obtained the lysis time by either method, what fraction of the lysis time did you allow for adsorption? Also, please add supplementary tables with lysis times used for the different phages.

      We thank the reviewer for this comment. We used published latent-period values as guides and verified compatibility with our own system when selecting incubation times. We have clarified this in the text and added the relevant citations. We did not use a common fixed fraction of the lysis time for all phages; instead, incubation times were chosen to allow sufficient time for adsorption but not for completion of the first lytic cycle. For λ, productive lytic development was blocked in the host background used, as in Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119. For ϕ80 and T5, we used published latent-period values as guides and verified their compatibility with our own system (De Paepe and Taddei, PLoS Biology 2006, http://dx.doi.org/10.1371/journal.pbio.0040193). M13 is a chronic filamentous phage and therefore does not have a standard lytic latent period; in our host–phage combination, it required more than 1 h before phage release. For T6, we relied primarily on the kinetics observed in our own system, since adsorption was unusually slow for this phage–host pair under our assay conditions. Although literature reports describe shorter T6 latent periods under specific assay conditions (Foster and Johnson, Journal of General Physiology 1951, http://dx.doi.org/10.1085/jgp.34.5.529), this is consistent with published work showing that adsorption and infection kinetics can vary substantially with host background, surface structure, and experimental conditions (Heller and Braun, Journal of Bacteriology 1979, http://dx.doi.org/10.1128/jb.139.1.32-38.1979; Storms et al., Biochemical Engineering Journal 2012, http://dx.doi.org/10.1016/j.bej.2012.02.010).

      (5) Supplementary

      Figure S1

      This data is useful in understanding the main body of the paper, and I think this should form part of a main figure (possibly with the individual experimental data points superimposed over the bars). This could come before or as part of Figure 1?

      We thank the reviewer for this suggestion. We have explored including these data directly in the main figure but found that doing so substantially reduced the readability of the figure, as the underlying table is visually dense. For this reason, we chose to summarize the results in Figure 1 and present the detailed data separately in Figure S1 of the Supplementary Material, along with the Ratio analysis, which more effectively conveys the trends without overloading the main figure.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) L16-18: This sentence could be made more accessible as 'error correction' is not an intuitive term in the phage field.

      We have updated the overall theory section including the terminology. Instead of error correction, we now refer to it as a discrimination process.

      (2) L96-98: Does this potentially indicate a trade-off where evolution for stronger binding cannot evolve at the same time as responsiveness to metabolic activity?

      We agree that this sentence made a stronger evolutionary claim than our data support. Since we only tested four laboratory phages, we cannot conclude that there is an evolutionary trade-off between stronger binding and responsiveness to host metabolic activity. We have therefore removed this sentence to avoid making an unsupported evolutionary interpretation.

      (3) L102: What does 'post-cellular' mean?

      Postcellular supernatant is simply the liquid that remains after cells have been removed. During centrifugation, the cells pellet at the bottom, and the liquid above (which can contain viruses) is the postcellular supernatant.

      (4) L105-107: Worth splitting into two sentences as it is a bit unclear if ratios are built between permissible and resistant hosts or between buffers or both.

      Thank you for the suggestion. We have rewritten this section into two sentences to clarify how the ratios are constructed, and we hope the revised wording improves readability.

      (5) L110-122: Figures S1 and S2 could be referenced here.

      References to Figures S1 and S2 have now been added in this section.

      (6) L137: As P(0) is the viral concentration in buffer, I am assuming that the phage lysate has been diluted in buffer and phages have been added to cultures from the same dilution tube to guarantee equal starting numbers, but I couldn't find this in the methods.

      This clarification has been added to the Methods and Media section of the Supplementary Information.

      (7) L243: It would be worth defining what 'hyperdiffusion' means.

      We have added a brief definition of “hyperdiffusion”.

      (8) L253-256: I do not entirely follow this explanation.

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #3. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (9) L284: Why is k_inj necessarily more sensitive to the metabolic state than k_off? Could membrane changes under stress increase k_off?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: reduced commitment in inactive cells can arise through an increase in k_off, a decrease in k_inj, or both, as long as k_off/k_inj becomes larger in the inactive case. What matters is therefore not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also underlies the trade-off now discussed in the manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become much larger in the inactive case. We have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (10) Figure 1: There seems to be more variation between replicates in phage Lambda than in other phages. Is this caused by receptor number heterogeneity in the population?

      Unfortunately we do not have a way to compare receptor number heterogeneity across the different phage receptors in our experiments. We therefore cannot conclude that the larger variation observed for phage λ is caused by receptor number heterogeneity in the population.

      (11) Figure S1: There seems to be a significant difference between phage Lambda viability in the two buffers - do the authors have an idea where this comes from?

      There is no difference in λ viability between the two buffers. The apparent difference in the figure is due to sampling variability.

      (12) Figure S3: Last sentence of the legend probably shouldn't say 'upper'.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      Reviewer #3 (Recommendations for the authors):

      (1) The text reads as incomplete in some places. Can the authors please provide clarifications on the following points?

      (1.1) Lines 235-256: How do the authors draw a conclusion that "a phage can detect host metabolic status" from a study that used purified LamB receptors (i.e., no live cells with any metabolism) extracted from two different bacterial species (i.e., not a difference in metabolic states)?

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #2. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (1.2) Line 270, in the abstract, and in the caption of Figure 4: The authors described the model using terms such as "an error-correction mechanism" or "standard error correction", but there is little explanation. Can the authors clarify what kind of "error" is discussed here, and how it is "corrected"? In the "standard error correction" model, what determines which method of correction is "standard"? If "error correction" is a standard term in phage-bacterial interaction modeling, please provide references.

      We agree with the reviewer that our use of the term error correction was not appropriate in this context. The proper term is discrimination process rather than error correction. We have now corrected this terminology throughout the manuscript and clarified the underlying logic in the relevant sections.

      (1.3) Line 301: The authors speculated that phage T5 is "better suited to ecological niches", but I am not sure how that is consistent with their data showing T5 is more rampant, that they infect both energy-competent and energy-depleted cells, not just depleted cells. Why "niches", and why are T5 better suited to environments "where energy-limited cells dominate", not just any environment?

      We agree that this point was not stated clearly enough. What we intended to convey is that T5 would be at a net disadvantage in a niche containing a mixture of energy-competent and energy-deficient hosts. We have updated the main text accordingly.

      (1.4) Line 303, and related to point 6.3. above: Phage λ can also infect and replicate in "starved bacterial cells" (shown in Kourilsky 1974 and Geng et al. 2024, both of which were cited in this manuscript). How do the authors reconcile these reports with the discussion point in line 303, and their data that only phage T5, but not λ, shows insensitivity to the host metabolic state?

      Our data do not imply that phage λ is unable to infect starved bacteria. As shown in Kourilsky (1974) and Geng et al. (2024), λ can indeed infect and replicate in nutrient-limited cells. Our results specifically indicate that λ infection under starvation proceeds with a reduced adsorption rate, while T5 maintains the same adsorption rate even when the host is starved. Thus, our conclusion is that T5 is insensitive to the host metabolic state at the level of adsorption, whereas λ is not. We acknowledge that the wording in line 303 may have unintentionally led to confusion, and we have revised this part of the text to avoid that.

      (2) The following comments relate to the text and figures in the manuscript. There are many places in the manuscript that could use fine proofreading and copy-editing for clarity and consistency. For example:

      (2.1) If I understand it correctly, the equation in between lines 109 and 110 should be clarified using terms such as "Free viral particles after mixing with bacteria in Arsenate and Azide" and "Free viral particles in bacteria-free buffer with Arsenate and Azide". As it stands, it is not clear which terms correspond to conditions where bacteria are present.

      The equation has been updated to explicitly indicate which terms refer to mixtures containing bacteria and which refer to bacteria-free controls, so that the correspondence between conditions is now clear.

      (2.2) Equations in between line 276 and 283, and elsewhere: Some concentration terms are enclosed in brackets ("[BP]"), while most are not.

      This notation has been clarified. We now use “[PB]” specifically to denote the transient phage–bacterium complex, distinguishing it from the product P⋅B. All other concentration terms are written without brackets for consistency.

      (2.3) Figure 4 and in equations: "BP" or "PB"?

      The notation has been made consistent throughout; we now use “PB” exclusively to denote the phage–bacterium complex.

      (2.4) Line 284 and line 286: The "inj" in "k_inj" is sometimes italicized, sometimes not.

      The notation has been standardized so that k_inj is now formatted consistently throughout the manuscript, without italicizing “inj.” Also we have replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (2.5) Figure 5: Was the step "Estimate cell numbers from OD" really performed on the next day after the experiment (i.e., >12 hours after infection and phage plating), not immediately after cell washing?

      Thank you for pointing this out. The figure has been updated to reflect the correct order of steps: cell numbers are estimated from OD immediately after washing, followed by overnight incubation of the plates.

      (2.6) Figure S1: As it stands now, the x-axis of each panel can be read either as "Permissive, Resistant bacteria, Buffer" (missing "bacteria" for the first pair of bars), or "Permissive (bacteria), Resistant (bacteria), Buffer (bacteria)" (extra "bacteria" for the last pair of bars).

      The intended interpretation is the second one (permissive bacteria, resistant bacteria, buffer).

      (2.7) Figure S3: The panel letters "A" and "B" are missing in the figure. Also, it is not clear why the legend for the five phages and the legend for the measurement times are not combined.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      (2.8) Strain table in the Methods and Materials: Please write genotypes with italicization, and consistently indicate mutations and deletions with the minus sign superscript or the Δ prefix. Also, for the S3222 strain: Is it really the entire Mal regulon mutated ("Mal-"), or just lamB-? In Brown et al. 2022, it was only the latter.

      Genotypes have been reformatted with consistent notation. For S3222, the correct designation is Mal-, as in the SI of Brown et al. 2022. In this case, Mal- is intended as a phenotypic designation rather than a specific genotype, and we have therefore formatted it accordingly, i.e. neither italicized nor written in lower case.

    1. eLife Assessment

      This study makes a valuable contribution by broadening the range of eukaryotic model systems and establishing Blastocystis, the most prevalent microeukaryote in the human gut, tractable for reverse-genetics investigations. The presented imaging data are convincing and informative, although confirmation by molecular methods would further strengthen the study. The work should interest readers studying host-microbe interactions in the human gut, as well as those developing new systems for eukaryotic research.

    2. Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in Blastocystsis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did - so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      Comments on revised version.

      The authors have revised their manuscript to clarify that molecular analyses have not yet occurred and have resolved the technical/publication issues with the figures. I look forward to seeing these tools used in future publications to answer important questions in Blastocystsis research.

    3. Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide a solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to visual presentation of the data would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested "23 promoter-terminator pairs" to better reflect that only promoters varied.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, the part A of Figure 2 is split into two panels with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of the part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible and it draws attention to a distracting design feature rather to the data themselves.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      (4.2) In the parts B-D, the legend should explain more clearly what each image shows and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images as the current distortions complicate interpretation.

      Comments on revised version.

      The revised version provides sufficient clarity and appropriate visual presentation. Some confusion evidently arose due to my misunderstanding, so I thank the authors for their comprehensive clarifications and patience.

    4. Author response:

      The following is the authors’ response to the original reviews.

      In revising the manuscript, we have focused on three main priorities raised during review: (1) improving precision around evidential claims, particularly concerning vector maintenance and P2A-mediated protein separation; (2) substantially improving figure quality, accessibility, and legend clarity; and (3) correcting inconsistencies and expanding methodological detail where requested.

      This study was intended as a foundational genetic toolkit and methodological framework for Blastocystis ST7-B, establishing practical workflows for DNA delivery, endogenous regulatory-element benchmarking, antibiotic-selected recovery, clonal propagation, and reporter-based analysis in a genetically challenging anaerobic microbial eukaryote. The central evidence presented is therefore functional in nature: reproducible transgene delivery, selectable recovery and propagation of colony-derived transgenic lines, and detectable reporter expression using multiple anaerobic-compatible reporter systems.

      We agree with the reviewers that several additional experiments, including Western blot analysis of P2A-containing constructs, outward-facing PCR, plasmid rescue assays, and selection-withdrawal experiments, would further strengthen the mechanistic interpretation of the system and help distinguish episomal persistence from genomic integration. We have therefore revised the manuscript throughout to clearly separate what is directly demonstrated from what remains a plausible working interpretation or important future direction.

      Importantly, the revised manuscript no longer presents episomal maintenance or complete P2A-mediated protein separation as demonstrated conclusions. Instead, these are now discussed explicitly as unresolved mechanistic questions requiring future molecular analysis. Nevertheless, the central methodological conclusion remains unchanged: stable selectable transgene expression, recovery of colony-derived transgenic lines, and reporter-positive Blastocystis ST7-B transformants can now be reproducibly obtained.

      Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in the Blastocystis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did, so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      We thank the reviewer for the positive assessment of the manuscript and for recognising both the technical difficulty and broader utility of establishing genetic tools in Blastocystis and other experimentally challenging microbial eukaryotes. We also appreciate the reviewer’s identification of the main evidential limitations in the original manuscript, particularly regarding vector maintenance and P2A-mediated protein separation.

      First, regarding the question of vector topology and the interpretation of episomal maintenance.

      We agree that the original manuscript presented episomal persistence too strongly relative to the evidence currently available. We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working model rather than a directly demonstrated conclusion.

      The transfection system used here was adapted from Li et al. (2019), including use of the pXS2-P<sub>Legumain</sub>-derived plasmid framework. Importantly, the construct used in the present study does not contain the original Trypanosoma brucei tubulin-targeting region associated with homologous integration in the original pXS2 system. Complete plasmid sequencing confirmed that the constructs function here as heterologous expression plasmids carrying Blastocystis ST7-B regulatory elements and transgenes. While this does not demonstrate episomal persistence, it also means that genomic integration cannot be inferred from the historical pXS2 vector architecture alone.

      We further note that comparative genomic analyses by Gentekaki et al. (2017) suggest that Blastocystis lacks components of the canonical non-homologous end-joining (NHEJ) machinery, implying that homologous recombination is likely to represent the principal route for double-stranded DNA repair. Because the constructs used here did not contain Blastocystis homology arms, there is currently no obvious mechanism favouring targeted homologous integration. Nevertheless, we fully agree that genomic integration cannot presently be excluded.

      To reflect this appropriately, the revised manuscript now explicitly separates the demonstrated functional outcomes from unresolved mechanistic questions concerning vector maintenance. We also identify several future approaches that would help distinguish episomal persistence from genomic integration, including outward-facing PCR, plasmid rescue followed by full plasmid sequencing, Southern blotting, FISH, selection-withdrawal experiments, and long-read sequencing approaches.

      We have revised the manuscript throughout to remove statements implying demonstrated episomal maintenance and now present episomal persistence only as a plausible working interpretation.

      In the Methods section under Cloning, the following text has been added:

      Lines 202–206: “The constructs used in this study were derived from the pXS2-P<sub>Legumain</sub> vector described by Li et al. (2019), which adapted a heterologous expression-vector backbone for transient plasmid-based expression in Blastocystis ST7-B. Here, the same molecular backbone was used as a plasmid scaffold carrying Blastocystis-derived regulatory elements and transgenes.”

      In the Discussion, the following text has been added/edited:

      Lines 665–673: “The molecular maintenance state of the introduced constructs remains unresolved: episomal maintenance is a plausible working model, but genomic integration cannot be formally excluded. The constructs used here lack Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components (Gentekaki et al., 2017), making targeted integration by standard repair routes unlikely but not impossible. Direct assays such as outward-facing PCR, plasmid rescue followed by full plasmid sequencing, FISH, or selection-withdrawal experiments will be required to distinguish episomal persistence from integration.”

      Second, regarding P2A-mediated protein separation.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. We have therefore revised the manuscript to avoid implying that complete protein-level separation was directly demonstrated.

      The revised manuscript now states only what is directly supported by the current data: that P2A-containing bicistronic constructs supported antibiotic-selected recovery of transgenic lines together with detectable downstream reporter expression. The microscopy data therefore support functional downstream reporter expression, but do not by themselves exclude residual uncleaved fusion products.

      We selected P2A because it is a compact and well-characterised peptide with high reported separation efficiency across multiple eukaryotic systems, including microbial eukaryotes. However, we agree that P2A performance can be context-dependent, and we now explicitly identify biochemical validation of P2A cleavage efficiency as an important future direction.

      Importantly, these revisions do not alter the central methodological conclusion of the study, namely that selectable transgene expression, propagation of reporter-positive lines, and recovery of colony-derived Blastocystis ST7-B transformants can now be reproducibly achieved.

      Text inserted in the Results:

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.”

      Lines 403–404: “However, protein-level separation was not directly tested, and the extent of any residual uncleaved fusion product remains unresolved.”

      Text inserted in the Discussion:

      Lines 619–629: “P2A was selected because it is a well-characterised peptide with high reported separation efficiency in human cell lines, zebrafish embryos, and mice (Kim et al., 2011). It also has precedent across microbial eukaryotes, including the protest Dictyostelium discoideum (Zhu et al., 2023), the fungi Aspergillus niger (Schuetze and Meyer, 2017) and Ustilago maydis (Müntjes et al., 2020), and the apicomplexan parasites Toxoplasma gondii (Markus et al., 2019) and Plasmodium falciparum (Dans et al., 2024). However, P2A performance is context-dependent, and the evidence presented here is functional rather than biochemical. P2A-containing constructs support antibiotic-selected recovery and downstream reporter expression in Blastocystis ST7-B, but ribosomal skipping efficiency and any residual uncleaved product will require direct protein-level validation.”

      Reviewer #1 (Recommendations for the authors):

      (1) Please could you show a Western blot to confirm if P2A has worked? It could be that the proteins are being expressed as a polyprotein.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. This is an important point, and we have revised the manuscript accordingly to avoid implying that complete protein-level separation was directly demonstrated.

      The current study was designed as a first-generation functional genetic toolkit for Blastocystis ST7-B, focused primarily on establishing reproducible workflows for selectable transgene expression, reporter recovery, and propagation of transgenic lines in this experimentally challenging anaerobic microbial eukaryote. The toolkit is therefore validated here through functional outcomes, including antibiotic-selected survival, stable propagation through extended passaging (>15 passages) and cryopreservation, and detectable reporter fluorescence above wild-type autofluorescence.

      P2A was selected because it is a compact and well-characterised peptide with high reported ribosomal skipping efficiency across multiple eukaryotic systems, including microbial eukaryotes, as discussed above. Nevertheless, we fully agree that direct biochemical validation would strengthen the mechanistic interpretation of the bicistronic system in Blastocystis ST7-B. We therefore now explicitly identify Western blot analysis, ideally using epitope-tagged upstream and downstream products, as an important future direction for quantitative assessment of P2A cleavage efficiency and any residual uncleaved fusion products.

      Relevant manuscript revisions are described above under the general response to Reviewer 1.

      (2) Something has gone wrong with figure formatting. Figure 2 is nearly illegible and I cannot read the text in section A. Sections B, C, and D have lost their labels and are fuzzy and surrounded by black. A similar issue affects Figure 3. Everything is just black with a few cells. It is illegible when printed.

      We thank the reviewer for highlighting these presentation issues and agree that the submitted figure quality significantly impaired readability and interpretation. The problems appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, contrast, and panel labelling in the review PDF.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved typography, panel separation, colour scaling, and legend clarity throughout. In response to additional reviewer suggestions, individual data points have now been added to Figures 2B and 2C to improve transparency and interpretability of the underlying data distributions.

      Figures 2 and 3 have been replaced with fully revised high-resolution versions with improved panel labelling, accessibility, typography, and figure legends.

      (3) The data from Figure 2B would be better placed in Table 1 with a column for robust/moderate/intermediate/weak/very weak. This would be much easier for the reader.

      We thank the reviewer for this helpful suggestion. We believe the comment refers to the promoter activity data shown in Figure 2A rather than the voltage optimisation data in Figure 2B. To improve readability and accessibility of these data, we have revised Figure 2A extensively to make the promoter activity tiers more legible and easier to interpret directly from the heat map and accompanying box plots.

      We considered incorporating simplified activity classifications into Table 1. However, activity patterns were construct-specific rather than simply locus-specific. In several cases, multiple promoter fragments derived from the same locus produced substantially different reporter outputs, and activity did not scale monotonically with promoter fragment length. We therefore felt that assigning a single categorical activity label at the locus level would oversimplify the dataset and reduce the construct-level resolution that is central to the toolkit value of the study.

      Instead, we addressed the reviewer’s concern by substantially improving the presentation and readability of Figure 2A, allowing readers to identify robust, moderate, intermediate, weak, and very weak expression constructs more directly while preserving the underlying construct-specific information.

      Figure 2A has been revised to improve clarity, accessibility, and legibility of the promoter activity tiers, allowing construct-level expression classes to be interpreted more directly from the heat map and accompanying boxplots.

      (4) How do you know if the constructs are maintained as episomes? Have you done an outward-facing PCR?

      We agree that direct molecular evidence distinguishing episomal persistence from genomic integration is currently lacking, and we appreciate the reviewer highlighting this important limitation. We have therefore revised the manuscript throughout to avoid presenting episomal maintenance as a demonstrated conclusion and now describe it only as a plausible working interpretation based on the current evidence and vector design.

      We have not performed outward-facing PCR in the present study. As discussed in the general response above, we now explicitly identify outward-facing PCR, plasmid rescue followed by full plasmid sequencing, selection-withdrawal assays, FISH, and long-read sequencing approaches as important future directions for resolving the molecular maintenance state of the constructs.

      The revised manuscript now clearly separates the demonstrated functional outcomes, including selectable transgene expression, recovery of colony-derived transgenic lines, and stable reporter-positive propagation under selection, from the unresolved mechanistic question of vector topology.

      This issue has been addressed throughout the revised manuscript, including in the Methods and Discussion sections, where episomal maintenance is now presented as a plausible but unconfirmed interpretation rather than a demonstrated conclusion.

      Minor Comments

      Line 66: is this one to two billion individuals with Blastocystis, or one to two billion Blastocystis cells per gut?

      The intended meaning was colonised individuals globally. We agree that the original phrasing was ambiguous and have corrected it for clarity.

      Lines 66–67 revised to: “…microorganisms in the human gut, and is estimated to colonise approximately one to two billion people globally (Scanlan and Stensvold, 2013).”

      Line 148: Supplier of IMDM?

      The supplier information was already present in the original manuscript as IMDM L0191 (Biowest).

      No additional manuscript change required.

      Line 157: Who annotated the dataset, the 2017 paper or the present study?

      The dataset annotation derives from Armengaud et al. (2017). We agree that the original wording was unclear and have revised this section substantially to improve clarity regarding the rationale and workflow used for promoter and terminator candidate selection.

      “The relevant Methods section has been extensively revised for clarity and expanded detail” (Lines 156–189).

      Line 166: Who predicted the 3′ UTR, the 2017 paper?

      This information derives from the NCBI annotation associated with the Blastocystis ST7-B genome based on Denoeud et al. (2011). This has now been clarified in the Methods section.

      Clarified in revised Methods section.

      Line 237: How long did it take in days?

      Approximately 2 days.

      Line 270 revised to: “…turned yellow without drug treatment, usually within 2 days post-transfection.”

      Line 325: Typo, missing gap between Figure and 1A.

      Corrected in revised manuscript.

      Reviewer #2 (Public review):

      This manuscript presents a substantial technical advance for the genetic manipulation of Blastocystis by establishing an integrated workflow for stable episomal transgenesis, antibiotic selection, clonal recovery, and reporter-based imaging in the ST7-B subtype. The study is particularly valuable because it combines multiple previously fragmented approaches into a coherent and practically applicable toolkit, including endogenous regulatory elements, optimized electroporation conditions, selectable markers, and anaerobic compatible fluorescent reporters. This methodological work greatly expands the molecular toolbox and future studies focused on both basic and infection biology can now build on the ability to express and localize proteins in fixed as well as live cells.

      The microscopy data are convincing and clearly demonstrate functional reporter expression and successful recovery of stable transgenic lines. Nevertheless, because this is primarily a methodological paper, the study would be further strengthened by the inclusion of Western blot validation of reporter expression and bicistronic constructs. In particular, biochemical analysis of the P2A-containing constructs would help assess the efficiency of ribosomal skipping and exclude the possible presence of uncleaved fusion proteins, thereby providing stronger support for the interpretation of the imaging data and the functionality of the expression system.

      We thank the reviewer for this thoughtful and positive assessment of the manuscript and for recognising the value of integrating previously fragmented approaches into a coherent and practically usable genetic toolkit for Blastocystis ST7-B. We particularly appreciate the reviewer’s recognition that the system expands the currently available molecular toolbox for both cell biological and infection-related studies in this experimentally challenging anaerobic microbial eukaryote.

      We also appreciate the reviewer’s comments regarding biochemical validation of the P2A-containing bicistronic constructs. We agree that Western blot analysis would strengthen the mechanistic interpretation of the reporter system by directly assessing ribosomal skipping efficiency and the possible presence of residual uncleaved fusion products. In response, we have revised the manuscript throughout to ensure that the conclusions remain appropriately evidence-based and do not imply that complete protein-level separation was directly demonstrated.

      The revised manuscript now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, stable propagation of reporter-positive lines, and detectable downstream reporter expression, from unresolved mechanistic questions concerning P2A cleavage efficiency and vector maintenance state. We now also identify biochemical validation of P2A-mediated protein separation as an important future direction for further refinement of the system.

      Relevant manuscript revisions addressing these points are described above under the response to Reviewer 1.

      Reviewer #2 (Recommendations for the authors):

      The quality of images could be better. The figures lacked resolution — possibly a conversion artefact.

      We agree that the figure quality in the submitted review PDF significantly reduced readability and visual interpretation. The issues appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, typography, panel labelling, and contrast rendering.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved panel separation, typography, colour scaling, contrast settings, and figure legends to improve accessibility and interpretability both on screen and in print. In addition, the export workflow and file formatting have been updated to improve compatibility with journal production requirements and reduce the likelihood of compression-related rendering artefacts during manuscript compilation.

      Figures 2 and 3 have been replaced with revised high-resolution versions with improved typography, panel labelling, contrast settings, and accessibility.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for the propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design, to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to the visual presentation of the data, would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      We thank the reviewer for this important point and agree that the original manuscript presented episomal persistence too strongly relative to the currently available evidence. In particular, we agree that the title and several sections of the manuscript implied a level of mechanistic certainty that was not directly demonstrated.

      We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working interpretation rather than a demonstrated conclusion. The revised text now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, recovery and propagation of colony-derived transgenic lines, and stable reporter-positive maintenance under selection, from the unresolved mechanistic question of vector topology.

      As discussed in our response to Reviewer 1, the constructs used here do not contain Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components, making targeted integration by standard repair routes less strongly supported mechanistically, although genomic integration cannot presently be excluded.

      We agree that selection-withdrawal experiments would provide useful additional evidence regarding construct persistence and have now explicitly identified such assays, together with outward-facing PCR, plasmid rescue, FISH, and long-read sequencing approaches, as important future directions for resolving the molecular maintenance state of the transgenes.

      The manuscript has been revised throughout to remove wording implying demonstrated episomal maintenance. Episomal persistence is now discussed only as a plausible working interpretation pending direct molecular validation.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      We thank the reviewer for this careful reading and for identifying this inconsistency. We agree that the distinction between candidate loci and successfully generated promoter constructs was not sufficiently clear in the original manuscript and could create uncertainty regarding the dataset.

      The original candidate set comprised 14 loci selected for promoter and terminator discovery. However, only 11 loci yielded successfully cloned and experimentally tested promoter constructs. The remaining three loci were retained in Table 1 for completeness and transparency, as repeated cloning attempts were unsuccessful despite two independent efforts.

      We have revised the manuscript to make this distinction explicit and to clarify that the reported NanoLuc benchmarking experiments were ultimately performed using constructs derived from 11 successfully cloned loci.

      Lines 361–364: “To expand the available regulatory parts, we screened 23 NanoLuc reporter constructs containing putative endogenous promoter–terminator pairs from 11 of 14 candidate loci; three loci could not be cloned after two independent attempts and are indicated in Table 1.”

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested “23 promoter-terminator pairs” to better reflect that only promoters varied.

      We thank the reviewer for the opportunity to clarify this point. We agree that the original presentation may have created the impression that promoter and terminator regions were independently varied and benchmarked, whereas the experimental design was primarily focused on construct-level comparison of endogenous regulatory modules.

      As described in the Methods, each construct contained a candidate endogenous upstream promoter region together with the corresponding endogenous downstream terminator region derived from the same locus. For consistency and to keep the cloning and screening strategy experimentally tractable, a fixed 500 bp downstream terminator fragment was used for each locus rather than systematically varying terminator length or independently testing terminator activity.

      We therefore retain the description “endogenous promoter–terminator pairs,” since each construct contains both endogenous upstream and downstream regulatory regions from the same genomic locus. However, we agree that the assay was not designed to independently dissect promoter versus terminator contributions to reporter output. We have revised the manuscript accordingly to make this distinction explicit and avoid ambiguity regarding the scope of the benchmarking analysis.

      Lines 365–368: “Each construct paired a candidate upstream promoter region with the corresponding downstream terminator region from the same locus, defined here as the native 500 bp sequence immediately downstream of the stop codon. Where multiple promoter lengths were tested for the same locus, the terminator fragment was kept constant (Table 1; Figure 1A).”

      This design allowed construct-level benchmarking of paired promoter–terminator modules but did not test promoter strength or terminator activity independently.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      We thank the reviewer for these important methodological points. We agree that the original manuscript did not sufficiently clarify the transient nature of the NanoLuc benchmarking assay or the rationale underlying the assay design and timing.

      The promoter benchmarking assay was designed as an early transient-expression screen adapted from the NanoLuc-based workflow of Li et al. (2019), with modifications, rather than as a stable-maintenance assay. No selectable marker was included because the objective was to compare relative early reporter output across constructs shortly after DNA delivery, before prolonged culture effects became dominant.

      The 16–18 h post-electroporation time point was selected based on the NanoLuc expression kinetics reported by Li et al. (2019) and empirical optimisation during assay development. This window allowed robust transient reporter detection while limiting confounding effects arising from prolonged plasmid loss, differential outgrowth, variable recovery, or later culture-level changes.

      We agree that, in the absence of selection, the observed NanoLuc signal cannot be interpreted as an absolute measure of promoter strength independent of DNA uptake efficiency, early plasmid retention, or post-transfection recovery dynamics. We have therefore revised the manuscript to clarify that Figure 2A reports relative transient reporter output under standardized early post-transfection conditions rather than isolated promoter activity alone.

      We now also explicitly state that the data shown in Figure 2A derive from three independent electroporation experiments per construct, each assayed in technical duplicate.

      Lines 241–248: “Promoter–terminator activity was assessed 16–18 h after electroporation using a transient NanoLuc assay adapted from Li et al. (2019), with modifications. This early time point was selected to capture reporter output within the transient-expression window after DNA delivery, before prolonged plasmid loss, differential outgrowth, or culture-level changes could dominate the readout. Because the constructs did not contain a selectable marker, the measured NanoLuc signal reflects early transient reporter output rather than promoter strength independent of DNA uptake, early plasmid retention, or post-transfection recovery.”

      Additional clarification added to Figure 2 legend stating that measurements derive from three independent electroporation experiments, each assayed in technical duplicate.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, part A of Figure 2 is split into two panels, with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible, and it draws attention to a distracting design feature rather than the data themselves.

      We thank the reviewer for these detailed comments regarding figure design and visual interpretation. We agree that the original presentation of Figure 2 introduced unnecessary visual ambiguity through inconsistent colour usage, the use of a diverging colour scale for strictly positive values, and poor readability associated with the dark background and low-resolution export.

      In response, Figure 2 has been extensively redesigned to improve clarity, accessibility, and interpretability. The previous blue–white–red diverging heatmap has been replaced with a sequential colour palette appropriate for positive-only expression data. Boxplot colouring has also been simplified and harmonised with the heatmap scheme to avoid implying unsupported categorical distinctions. In addition, panel organisation, typography, scaling, and legend structure have all been revised to improve readability and reduce ambiguity regarding interpretation of the plotted values.

      We also agree that the black background distracted from the data presentation and impaired visibility of panel labels and image boundaries. The revised figures therefore use white backgrounds together with clearer panel separation and improved label visibility throughout.

      Figure 2 has been completely reformatted using a sequential colour scale in panel A, simplified and harmonised boxplot colouring, larger typography, improved panel separation, revised legends, and white backgrounds throughout. Corrected high-resolution source figures have been provided.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      We thank the reviewer for this helpful comment regarding figure composition and panel separation. We agree that the original presentation made it difficult to distinguish intentional panel boundaries from imaging artefacts, particularly in the low-resolution review PDF generated during manuscript compilation.

      To improve clarity, Figure 3 has been reformatted using white backgrounds, clearer panel spacing, and more explicit separation between individual snapshots and imaging channels. High-resolution source images have also been provided to ensure that fluorescence patterns, image boundaries, and panel organisation remain clearly interpretable both on screen and in print.

      Figure 3 has been reformatted with improved panel separation, white backgrounds, clearer image boundaries, and revised high-resolution source figures.

      (4.2) In parts B-D, the legend should explain more clearly what each image shows, and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      We thank the reviewer for these helpful suggestions regarding figure annotation and legend clarity. We agree that the original presentation did not sufficiently explain the composition of the imaging panels, particularly under the low-resolution conditions of the review PDF.

      To improve interpretability, the revised Figure 3 now includes clearer panel organisation, improved annotations, and expanded figure legends explicitly identifying the individual imaging channels and staining conditions shown in each subpanel. The leftmost panels in parts B–D are now more clearly identified in both the figure and legend, together with the corresponding fluorescence or staining conditions used in each experiment.

      As mentioned in the Methods sections we used Hoechst 33342 to visualise DNA; but we agree that the Hoechst 33342-labelled structures required additional clarification. The revised legend section now explains that multiple Hoechst 33342-positive regions are commonly observed because Blastocystis cells can contain multiple nuclei depending on cell stage and subtype-specific morphology.

      In addition, high-resolution source images have been provided to ensure that fluorescent signals, panel boundaries, and imaging features remain clearly interpretable both on screen and in print.

      Figure 3 legends and annotations have been revised to clarify imaging channels, staining conditions, and panel organisation. The figure caption was also edited to include: “DNA was visualised using Hoechst 33342. Most cells contained two nuclei, and smaller Hoechst 33342-positive signals consistent with mitochondrial DNA were also observed in some instances.”

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images, as the current distortions complicate interpretation.

      We thank the reviewer for this important observation and agree that the apparent morphological differences between the UnaG/smURFP and SNAP-tag panels required additional clarification.

      The images shown for the different reporter systems were acquired under different imaging conditions and microscope configurations following an instrument-related issue during part of the imaging workflow, as noted in the Methods section. As a result, direct visual comparison of cell morphology between reporter systems is not appropriate. The primary purpose of these panels is instead to demonstrate reporter detectability, live-cell labelling capability, and the characteristic fluorescence patterns obtained with the different anaerobiosis-compatible reporter systems.

      In particular, the SNAP-tag panels were included to demonstrate successful live-cell labelling without permeabilisation together with the expected increase in fluorescence signal at higher substrate concentrations, rather than to support quantitative comparison of cell morphology across imaging conditions.

      We considered replacing the affected images. However, equivalent replacement datasets acquired under directly comparable conditions are not currently available. We have therefore retained the original images but revised the figure legend to clarify the intended interpretation and limitations of these panels explicitly.

      Figure 3 legend revised to include:

      “Because images for the different reporter systems were acquired under different imaging conditions, they are presented to demonstrate reporter detectability and labelling pattern and should not be used for quantitative comparison of cell morphology across reporter systems.”

      Reviewer #3 (Recommendations for the authors):

      The reader may find the current order confusing starting with construct design before testing which drug to use for selection. The narrative would work better if it started with antibiotic selection as the first logical step for generating stable cell lines.

      We thank the reviewer for this thoughtful suggestion regarding narrative structure and agree that multiple organisational strategies are possible for presenting a methodological workflow of this type.

      We considered reorganising the Results section to begin with antibiotic selection and drug sensitivity profiling. However, we ultimately retained the overall structure because the manuscript is organised as a toolkit-development framework rather than as a strictly chronological experimental protocol. The Results therefore begin with regulatory-element discovery and construct design, which form the conceptual and experimental foundation of the toolkit, before progressing to DNA delivery optimisation, drug sensitivity profiling, clonal recovery, and reporter validation.

      We felt that this structure most clearly reflects the dependency relationships within the system: regulatory elements are required before constructs can be assembled, constructs are required before electroporation conditions can be evaluated, and selectable constructs are required before stable selection and clonal recovery can be meaningfully assessed.

      (2) The text states that the screen 'focused on the 1,000 most abundant proteins to establish a preliminary library capable of supporting varying levels of transcription.' Since the genome has ~6,000 protein-coding genes, the top 1,000 cover the most abundant proteins — not a wide expression range.

      We thank the reviewer for this important clarification. We agree that the original wording could incorrectly imply that the screen was intended to sample broadly across the full transcriptional range of the Blastocystis genome. This was not the case, and we have revised the manuscript accordingly.

      Our strategy was instead designed to enrich for candidate loci with a higher prior likelihood of supporting detectable transgene expression. Because no genome-wide promoter map, transcription start site dataset, or experimentally validated regulatory annotation was available for Blastocystis ST7-B at the inception of this work, we used the abundance-ranked Blastocystis ST4-WR1 proteomic dataset of Armengaud et al. (2017) as a practical starting point for candidate discovery.

      Importantly, the Blastocystis ST4-WR1 proteome is highly skewed, with 193 proteins contributing approximately 50% of the detected proteome and the 13 most abundant proteins contributing approximately 10% (Armengaud et al., 2017). We therefore selected the top 1,000 proteins not as a representation of the genome-wide expression range, but as a proteomics-guided enrichment strategy to identify loci more likely to contain active endogenous regulatory regions suitable for initial toolkit development.

      We have revised the relevant Methods section substantially to clarify both the rationale and the workflow used for candidate selection, homolog identification, and promoter/terminator definition.

      The Methods section (Lines 156–189) has been extensively revised to clarify the rationale underlying candidate regulatory-element selection. The revised text now explicitly states that the strategy was designed to enrich for likely active loci for toolkit development rather than to systematically survey the full range of promoter strengths across the Blastocystis genome.

      Additional methodological detail has also been added regarding:

      Use of the Armengaud et al. (2017) proteomic and proteogenomic datasets,

      Homolog identification in Blastocystis ST7-B,

      Locus selection criteria,

      Promoter boundary definition,

      And operational definition of candidate terminator regions.

      (3) The Methods contain an inconsistency: cells were left in 0.5 mL, then 1 mL was added, but then only 0.5 mL is apparently used for transfection. What happened to the 1 mL?

      We thank the reviewer for identifying this ambiguity in the transfection workflow description. The apparent inconsistency arose because the protocol description moved from bulk cell resuspension to preparation of individual electroporation reactions without explicitly stating how the intermediate suspension was used.

      After washing, approximately 0.5 mL of cytomix buffer remained above the pellet, and 1 mL of complete cytomix buffer was then added to generate an approximately 1.5 mL cell suspension. Cells were counted from this pooled suspension, after which the volume corresponding to 5 × 10<sup>7</sup> cells was transferred into each individual electroporation reaction. Following addition of DNA, each electroporation reaction was adjusted to a final volume of 500 µL with complete cytomix buffer. The remaining cell suspension was retained for additional transfections or control reactions.

      We agree that the original wording could be misinterpreted and have revised the Methods section to clarify the sequential handling steps more explicitly.

      Lines 225-229 revised to read: “The resulting approximately 1.5 mL pooled cell suspension was used for total viable cell counting using a hemacytometer.”

      “After counting, the volume corresponding to 5 x 10<sup>7</sup> cells was transferred to each electroporation reaction and combined with 25 µg of plasmid DNA. The total electroporation volume was adjusted to 500 µL with complete cytomix buffer.”

      (4) Figures 2 and 3 are too low-resolution for the font size used and for clearly viewing the microscopy images.

      We thank the reviewer for highlighting these readability issues. As noted in our responses above regarding Figures 2 and 3, the low-resolution appearance primarily resulted from manuscript compilation and PDF export artefacts affecting typography, image rendering, and panel clarity in the review version.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions featuring improved typography, panel labelling, contrast, accessibility, and image clarity for both on-screen viewing and print reproduction.

      Revised high-resolution versions of Figures 2 and 3 have been provided as described above. No additional manuscript changes were required beyond the figure revisions already outlined.

      (5) Figure 4 is confusing because the left and right panels appear inconsistent, with much higher concentrations required for growth inhibition in the culture-based assay than the resazurin assay indicated. The rationale for the resazurin assay should be explained, and the complete growth inhibition (CGI) concentration should be highlighted in the right panel.

      We thank the reviewer for highlighting this potential source of confusion. We agree that the distinction between the two assay endpoints was not sufficiently emphasised in the original figure presentation and legend.

      The apparent discrepancy arises because the two assays measure different biological endpoints under different assay conditions. The resazurin assay was used to estimate IC<sub>50</sub> values, corresponding to the concentration at which metabolic activity was reduced by approximately 50% under the assay conditions. In contrast, the small-culture assay was designed to determine complete growth inhibition (CGI), defined operationally as the concentration at which no detectable culture outgrowth occurred after incubation, using phenol red acidification as a culture-level readout.

      Because these assays measure partial metabolic inhibition versus complete suppression of detectable culture outgrowth, the corresponding concentration ranges are not expected to coincide directly. The higher concentrations observed in the right-hand panels therefore reflect the more stringent endpoint associated with complete growth inhibition rather than inconsistency between the assays.

      We agree that this distinction should have been explained more clearly in the original manuscript. We have therefore substantially revised the Figure 4 legend to clarify the rationale underlying both assays, explicitly distinguish IC<sub>50</sub> and CGI endpoints, and explain how the CGI values were used to guide subsequent antibiotic selection conditions for Blastocystis ST7-B transformants. The CGI transition range has also been made more visually explicit in the revised figure presentation.

      Figure 4 caption revised to: “Antibiotic potency and selection-window determination in Blastocystis ST7-B. Dose–response curves for puromycin, trimethoprim, and WR99210 were estimated from a resazurin-based viability assay (n = 3 independent replicates per drug per concentration). Points show mean ± SD, and the insets list the estimated IC50 values with R<sup>2</sup>-values > 0.75 for all fitted curves. The IC<sub>50</sub> estimates represent the drug concentrations that reduced resazurin-based metabolic activity by 50% under the assay conditions.”

      Right panels: “small-culture complete growth inhibition assay using 1 × 10<sup>7</sup> WT Blastocystis ST7-B cells per culture, assayed in triplicate across a wide range of concentrations. Cultures were incubated for 2 days, and outgrowth was assessed using phenol red acidification of the medium as a culture-level readout, with yellow indicating growth and red indicating no detectable growth. The yellow-to-red transition was used to estimate the concentration required for complete growth inhibition and to guide the subsequent antibiotic selection strategy for Blastocystis ST7-B transformants.”

      “IC<sub>50</sub> and CGI represent distinct assay endpoints: the former measures partial reduction in metabolic activity, whereas the latter identifies the concentration at which no detectable culture outgrowth occurs under the small-culture assay conditions.”

      (6) In Figure 3B, the unexpected UnaG fluorescence pattern could be due to protein sequestration because the protein is mildly toxic to the cell. This should be discussed in addition to the reasons already provided.

      We thank the reviewer for this thoughtful suggestion and agree that protein sequestration or reporter-associated cellular stress represent plausible alternative interpretations of the observed UnaG fluorescence pattern.

      We considered the possibility of UnaG-associated toxicity during interpretation of these data. However, under the conditions tested, we did not observe clear evidence of a substantial toxic effect: UnaG-expressing Blastocystis ST7-B cells could be recovered as stable lines, maintained under antibiotic selection, and propagated through continued culture. We therefore felt that direct attribution of the observed fluorescence pattern to reporter toxicity would currently remain speculative.

      At present, we consider the biochemical properties of the UnaG system itself to provide a more parsimonious explanation for the observed localisation pattern. In particular, unconjugated bilirubin is highly hydrophobic and would be expected to partition preferentially into lipid-rich cellular environments. This interpretation is consistent with the lipid-rich peripheral and intracellular structures previously reported in Blastocystis ST7-B (Liao et al., 2023).

      We have therefore revised the Discussion to acknowledge that the observed UnaG fluorescence pattern may reflect a combination of reporter-specific biochemical behaviour, bilirubin partitioning, local intracellular environment, or possible sequestration phenomena. At the same time, we avoid assigning toxicity as a demonstrated mechanism in the absence of direct measurements of cell fitness, reporter abundance, or bilirubin distribution. Such experiments would be required to evaluate this possibility rigorously.

      Lines 642-648: “Consistent with this, lipid-rich peripheral and intracellular structures have been reported in Blastocystis ST7-B, potentially providing favourable microenvironments for BR partitioning and contributing to the punctate UnaG fluorescence pattern (Liao et al., 2023). An alternative possibility is that the observed signal pattern reflects reporter sequestration or reporter-associated cellular stress. However, because UnaG-expressing lines were recovered, maintained under selection, and propagated through continued culture, toxicity remains a possible but untested explanation rather than a demonstrated mechanism.”

      Minor Comments

      Figure 2: Parts B and C should also show individual datapoints for better reader assessment.

      We agree that inclusion of individual data points improves transparency and interpretability of the underlying data distributions.

      Individual data points have now been overlaid on the boxplots in Figures 2B and 2C.

      Figure 3A: Separate channels (fluorescence, bright-field, merge) should be shown rather than only the merge. The current overlay is difficult to interpret, especially for colour-blind readers.

      We appreciate the reviewer’s concern regarding accessibility and interpretability. We considered separating the fluorescence, bright-field, and merged channels for Figure 3A. However, this panel was intended primarily as an overview demonstrating reporter detectability within the bicistronic construct context, while the detailed fluorescence distribution is explored more extensively in the subsequent UnaG panels. We therefore retained the merged presentation for Figure 3A. Importantly, the image is not dependent on red–green discrimination, as it combines a greyscale bright-field background with a high-contrast green/cyan fluorescence signal that remains distinguishable through brightness and contrast differences. In addition, colour-blind-friendly lookup tables (LUTs) were used throughout the revised figure set.

      To further improve accessibility, the original red annotation arrow has been replaced with a colour-blind-friendly annotation colour.

      Briefly define system components (P2A, UnaG, smURFP, SNAP-tag) and add an abbreviation list.

      We agree that brief contextual definitions improve accessibility for readers less familiar with these reporter systems. Rather than adding a separate abbreviation list, we have added short explanatory descriptions at the points where these components are first introduced in the manuscript.

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.” Line 515–516: “UnaG, a bilirubin-binding fluorescent protein originally isolated from the muscle of the Japanese eel (Kumagai et al., 2013)…” Line 527: “smURFP (small ultra-red fluorescent protein)…”

      Abstract: “among the most prevalent microbial eukaryote” should be “eukaryotes”.

      Corrected in revised manuscript.

      Conclusion (2nd sentence): unclear what “endogenous regulatory part discovery” means.

      We agree that this phrase required clarification. The intended meaning was the identification and benchmarking of native Blastocystis ST7-B promoter and terminator elements for construct design and toolkit development. We have clarified this directly in the revised Conclusion section.

      Lines 682–683 revised to: “By bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements, …”

      Author contributions: “critical advise” should be “advice”.

      Corrected in revised manuscript.

      Again, we thank the reviewers for their careful evaluation, constructive criticism, and thoughtful feedback on the manuscript. The review process has substantially strengthened the manuscript by helping us clarify the distinction between what is directly demonstrated experimentally and what remains mechanistically unresolved.

      The central methodological conclusions of the study remain unchanged: the toolkit enables selectable transgene expression, recovery of colony-derived lines, and propagation of reporter-positive transgenic Blastocystis ST7-B lines, extending genetic accessibility in this organism substantially beyond the previous transient transfection framework.

      At the same time, the revised manuscript now more explicitly acknowledges important unresolved mechanistic questions, including vector topology, P2A-mediated protein separation efficiency, and persistence in the absence of selection. These are now discussed transparently together with the future experimental approaches that will be required to address them directly.

      We believe the revised manuscript now presents a clearer, more rigorous, and more accessible description of a practical genetic toolkit for Blastocystis ST7-B and hope that the revisions and clarifications satisfactorily address the reviewers’ concerns.

    1. eLife Assessment

      This important study explores whether complex structures that are lost during evolution can re-evolve, which is a long-standing debate in evolutionary and developmental biology. The authors demonstrate that re-evolution can occur if the gene regulatory network that underlies the development of complex traits is maintained. The evidence supporting its conclusions is convincing and the work will be of interest to those studying the evolution and development of complex traits.

    2. Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype re-evolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      Comments on revised version.

      The authors have addressed the concerns in the revision. No further comments.

    3. Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species, and in different ant castes within these species. The Authors show that this is linked to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve as the key genes have earlier developmental roles, such that CRISPr knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      As the authors note in their response this is very difficult to achieve. While the addition of this data would raise this manuscript to an outstanding one, I think the data presented is solid, well-presented and provides novel insight.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Vasquez-Correa and colleagues describes the expression pattern of the ocelli (simple eye) gene regulatory network in ants. They correlate the expression pattern of these genes with the presence and absence of ocelli in different classes and species of ants. The presence of ocelli is a polyphenic trait in ants - understanding the molecular and developmental underpinnings of polyphenic traits is of significant interest to evolutionary biologists, developmental biologists, and ecologists. The authors propose that the presence of the latent expression of the ocellar network in classes of ants that do not display ocelli in the adults may underlie the re-evolution of ocelli within the ant lineage.

      Strengths:

      The strengths of the manuscript are that it is well written, the images are of the highest quality, and the data support the conclusions of the authors.

      We thank Reviewer 1 for their positive comments.

      Weaknesses:

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants. A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are, in fact, responsible for ocelli development in the ant.

      The reproductive caste in ants is typically composed of both winged males and winged queens. We agree with Reviewer 1 that the queen caste, which develop 3 fully functional ocelli, is an important point of comparison in our study to the wingless minor workers and soldiers. Unfortunately, however, laboratory colonies rarely produce reproductive queens, and in the field, queen production in colonies of C. floridanus occurs within a narrow seasonal window, making the collection of queen larvae particularly challenging for developmental work. In contrast, the winged males, which also develop 3 functional ocelli like the queens for help during mating flights, can be readily generated in the lab throughout the year. Therefore, we use males as a proxy for characterizing ocelli development and GRN in queens and the winged reproductive caste as a whole. Given the deeply conserved gene regulatory networks underlying this trait across insects, we believe this is a reasonable assumption.

      We also agree with Reviewer 1 that using RNAi to knock down genes in the ocelli GRN would improve the study. For completeness of the scientific record, we would like reviewers and readers to know that we actually did, in fact, invest significant effort trying to knock down otd-1 (ortholog of the Drosophila otd gene), which functions as key upstream regulator of ocellar development. In Drosophila, RNAi knockdown of otd disrupts the development of all three ocelli as well as fine morphological features on the anterior of the head. In C. floridanus, otd -1 is expressed in the head capsule and brain (see Author response image 1 in this response). Injection of dsRNA of otd-1into whole soldier-destined larvae, significantly reduced otd -1 expression in the brain relative to its control, while in the head capsule, otd -1 expression remained largely unchanged relative to its control (see Author response image 1 in this response). This indicates that in the same individual, the injected otd -1 dsRNA was able to penetrate and significantly reduce otd -1 expression in the brain, but, was unable to penetrate the head capsule, where otd -1 expression remained largely unchanged. No ocellar phenotypes could be observed in pupae or adults. Therefore, for technical (not biological) reasons, we were unable to knockdown genes in the ocelli GRN in the head capsule. We hope to solve this technical problem in the coming years to add a mechanistic explanation for the latent expression and maintenance of the ocelli GRN in workers that completely lack ocelli as adults.

      Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype revolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      We thank Reviewer 2 for their positive comments.

      Weaknesses:

      Although the molecular pathway is conserved, the mechanism underlying the lack of ocelli in workers remains unclear. In C. floridanus, it could be explained by the evidence of no expression of certain developmental genes, but in other species, e.g. Polyrachis rastellata, is their expression intact, or reduced? There is no control male.

      In addition, HCR in species with partial re-evolution (if their genomes have been sequenced) would be useful to understand the mechanism. For example, there might be differential spatial expression between medial and lateral ocelli.

      We agree with Reviewer 3 that investigating the mechanisms underlying the lack of specific ocelli in these and other species is the next step for this research. Here, our main focus was instead on trying to explain the mechanisms underlying partial reversion of ocelli through the persistence of ocelli GRN expression in adult workers lacking ocelli. We therefore focused on the latent expression of the ocelli GRN in Polyrachis rastellata, a species that completely lack ocelli in adult workers, and how it may have facilitated the partial reversion of a single ocellus in its congener Polyrachis bihamata. Therefore, although we did not reveal specific interruption points in the ocelli GRN in Polyrachis rastellata, our results showing that this species expresses three genes of the ocelli GRN, offers sufficient evidence that this network is conserved and likely facilitated the partial reversion to a single ocellus in P. bihamata.

      We also agree with Reviewer 3 regarding the male control in P. rastellata and obtaining the species in our study that have undergone partial re-evolution. Unfortunately, these ants occur in Southeast Asia and are very difficult to collect. For males in P. rastellata, our colony died before we could try to induce male development. However, given the deep conservation of the network in the males of a genus within the same subfamily (Camponotini), we feel it is reasonable to assume that the network would also be conserved in the males of P. rastellata, especially since the genes we sampled are conserved in workers that do not develop ocelli as adults. As am sure the Reviewer may know that this is a continual challenge of working with emerging models in evodevo.

      Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species and in different ant castes within these species. The authors show that this is linked to to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstatrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      We thank Reviewer 3 for their positive comments.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve, as the key genes have earlier developmental roles, such that CRISPR knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      We agree with Reviewer 3 that functional experiments in species where workers both have ocelli present and absent is a key next step in this research. Please see our response to Reviewer 1 on our failed attempts to achieve this. We are therefore grateful to Reviewer 3 for acknowledging the challenges in trying to establish RNAi and CRISPR in the head capsules of developing workers in these ants. Also, please see our response to Reviewer 2 on the difficulty of finding and collecting these ants, which occur mainly in Southeast Asia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants.

      A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are in fact responsible for ocelli development in the ant.

      Please see our response to Reviewer 1 above.

      Reviewer #2 (Recommendations for the authors):

      For the questions below, if there is no experimental evidence, consider addressing them in the Discussion.

      Do sizes of ocelli different between castes? For example, even workers have 1-3 ocelli, their sizes are smaller than those of males/queens, especially in workers with 1 ocellus. If so, might it be continuous (not binary) changes in downstream gene expression that control ocellus size, with no ocellus below threshold? Does this favor the hypothesis of threshold but not switch?

      We thank Reviewer 3 for highlighting an important point about the size and development of ocelli. Observations suggest that ocelli tend to be larger in queens and males than in workers and soldiers in species with ocelli. However, we lack quantitative data to test this conclusively. We now include a sentence on the Discussion stating that an important avenue of future work should investigate whether threshold or switch mechanisms influencing the presence/absence, as well as size, of ocelli between queens and workers.

      For the species whose workers have a single ocellus, are there variations, e.g. spanning from 0, 1 to 2? If 2, always one medial plus one of the two laterals? If always one, it would be a good control for staining to see up- vs down-regulation of downstream gene expression within the same individual.

      We agree with Reviewer 2 that this is a fascinating approach to our question. We have not observed natural wild-type variation in the number of developing ocelli in the same-sized individuals in the worker caste. However, in a distantly related leaf-cutting ant species (Atta cephalotes) belonging to different subfamily (the Myrmicinae) individuals with different head-to-body scaling within the same colony can vary in the number of ocelli. For example, soldiers of Atta cephalotes include individuals developing one, two, or three ocelli. These configurations can appear as only the median ocellus, only the two lateral ocelli, or even the median plus a single lateral ocellus. Interestingly, these correlations vary with changes in the size and head-to body scaling, suggesting that each ocellus can undergo different degrees of development, with one or more remaining vestigial or completely absent. On the other hand, workers in other species consistently develop a single ocellus, like in workers of Polyrachis bihamata, with no correlation to size or head-to-body scaling. These cases highlight how evolutionarily labile this trait is among workers of different ant species, which supports our proposal that the underlying gene regulatory network remains latent, thereby facilitating the emergence of novel trait combinations. We therefore agree on the importance of comparing the developmental mechanisms underlying these patterns temporally across larval stages and between individuals within a colony. We have now incorporated 2 sentences into the discussion, stating that this will be an important avenue for future work.

      Is there any function of a single ocellus in workers, or just a consequence of incomplete down-regulation of gene expression?

      Thank you again for highlighting these important points that help us to elaborate on the discussion of our study. The functional role of ocelli in species that develop these structures remains largely understudied. However, for some species particularly within the Formicinae clade the function of the three ocelli in workers has been investigated, revealing that they serve as a celestial compass that facilitates navigation. We reference these findings in our Introduction and Discussion to illustrate that the presence of three ocelli in workers can represent an adaptive trait. In contrast, the functional significance of a single ocellus or of partially developed ocelli remains an important question. This knowledge gap presents a promising avenue for future research to understand the adaptive value of reduced, partially suppressed ocellar development. We have now added a sentence in the discussion stating this.

      In previous studies, JH treatment can increase the number of ocelli in workers, consistent with its role in promoting reproductive development. In the ocellus developmental pathway, what causes the reduction of downstream gene expression in C. floridanus? Does JH directly regulate their expression?

      We thank Reviewer 2 for proposing yet another interesting question for future investigation, which we have added to the Discussion.

      The only current evidence available in C. floridanus is a recent study (MacMillan et al. 2025), in which minor workers and soldiers were treated with JH at different developmental stages. Unfortunately, no evidence of ocelli induction was observed in JH-treated individuals, suggesting that the mechanisms of ocelli development in C. floridanus might be highly canalized, especially in species that exhibit worker polymorphism (inter-individual variation in size and head-to-body scaling within the worker cate). However, more studies are required to understand why in Monomorium pharonis (no worker polymorphism) ocelli development can be readily induced by JH, while in another C. floridanus (with worker polymorphism) it appears quite difficult.

      "In D. melanogaster, the head develops from the eye-antenna disc" This statement is not correct. The brain does not belong to the eye-antennal disc.

      We thank Reviewer 2 for catching the misspelling. We have changed the name to eye-antenna disc in the sentence.

      Reviewer #3 (Recommendations for the authors):

      It is hard to see the developing ocelli in Figure 7 - could the authors increase the contrast to make them more visible?

      We have made the suggested changes to Figure 7 in the main article, and it has indeed improved the figure.

      Author response image 1.

      RNAi knockdowns in developing soldiers of Camponotus floridanus show a reduction of otd -1 expression in the brain, but no effect on otd -1 expression in the eye-antenna disc. A. HCR revealing otd -1 expression in the brain B. qPCR of otd -1 expression after RNAi knockdown shows significantly reduced otd -1 expression in the brain, C. HCR revealing otd -1 expression in the eye-antenna disc D. qPCR of otd -1 expression after RNAi knockdown shows no significant affect on otd -1 expression in the eye-antenna disc.

    1. eLife Assessment

      Building on earlier studies, this manuscript reports a role for pol kappa in cisplatin resistance in the very specific scenarios of head and neck squamous cell carcinoma, providing evidence that the PIP box of Pol kappa is critical for cisplatin resistance in these cells. The findings are of a highly focused relevance and will be useful in the field, but the conclusions are limited to very specific cancer cells. Conclusions cannot be generalized to all cisplatin resistance mechanisms and cell types and are based on incomplete evidence that presents uncertainties and discrepancies that need to be resolved.

    2. Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA-interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data. USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism. The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Overinterpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polk in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014). Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

    3. Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

    4. Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Thank you so much for appreciating our efforts to demonstrate role of Polκ mediated axes in cisplatin resistance in head and neck cancer cells.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      We completely agree with you, and the presented data only support Polκ's role in HNSCC as demonstrated in both acute cisplatin exposure as well as the cisplatin-resistant HNSCC models. Other cell and cancer types may adopt different strategies for cisplatin resistance.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      Since H357-S and SCC9-S cells are highly sensitive to cisplatin, knocking down of Polκ unlikely will alter the phenotype, as other TLS DNA polymerases like Polκ and Polκ play critical role in such lesion bypass. Since no other DNA polymerase was upregulated in these cells upon cisplatin exposure and in the cisplatin-resistant cells, it was intriguing to demonstrate a direct role of Polκ in chemoresistance and that has been proven in this study. Since the catalytic activity of Polκ is not required to induce chemoresistant in these cells, we strongly believe that Polκ-dependent mutagenesis play minimal or no role in adapting cells to tolerate cisplatin. Nevertheless, we will knock down Polκ in these cells and determine cisplatin sensitivity

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data.

      We appreciate the Reviewer's critical and insightful comment. In our view, the interaction between Polκ and USP18 is very specific as USP2 and IgG alone do not pull down Polκ. Similarly, we also show that both Polκ and USP18 interact with PCNA. We agree with the reviewer that the existence of a stable complex of PCNA-Polκ-USP18 has not been fully demonstrated in the current version. We will perform additional experiments to strengthen our finding: a) Co-IP experiments with Polκ PIP mutants (wild-type vs. mutant) should be performed to determine whether USP18 loses its ability to bind PCNA in the absence of Polκ-PCNA interaction. b) Mapping the domain in Polκ that is involved in USP18 binding and their Co-IP experiment. Additionally, high resolution co-localisation images including quantified data will be provided.

      USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism.

      We find the point raised by the reviewer is very intriguing, however, as it will require a significant amount of time and effort to demonstrate ISGylation of DDR proteins and deISGylation by UPS18, and the insight that we may gain is unlikely to add to the central theme of this paper, we will expand this in our subsequent related study. Thank you for the suggestion.

      The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Over interpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polκ in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014).

      Thank you very much for the suggestions. We will take care of the portions and modify as suggested. The necessary reference will be added as appropriate.

      Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

      Thank you for pointing this out and we will modify the portion for better clarity.

      Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Thank you for appreciating our finding that the PIP box of Polκ is critical for cisplatin response

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      We are extremely sorry for the lack of clarity. Please note that two cisplatin-resistant models (H357 and SSC9) have been used and the results were very consistent in both cells. Fig. 1C clearly demonstrates about the generation of these resistant models and the original reference has been already cited.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

      Thank you for pointing this out. In our view, the increased G2 arrest is due to more fork stalling or collapsed than the new origin firing. Also, in our assay we observed less than 10% of new origin fired DNA fibres, and that could be the reason of no significant change in new origin firing among various cells.

      Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      Thank you very much for finding our study interesting and the pending concerns will be addressed as suggested.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      In this study, we have explored eight different cell types (breast, brain, liver, head and neck, pancreatic, prostrate, lungs, and kidney) to check the expression of Polκ upon cisplatin exposure, and HNSCC cells only showed Polκ up-regulation. Therefore, we went ahead for further demonstration of the role of Polκ in cisplatin resistance in OSCC using four different cell models (H357-S, H357-R, SSC9-S, and SSC9-R). By adding more cell lines to study will unlikely change the central theme of the paper. Yes, by acquiring and analysing clinical samples from the cisplatin responder and non-responders would have strengthen our finding.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      It’s an interesting point and we will check whether overexpression of Polκ in H357-S cells could induce resistance to cisplatin and alters IC<sub>50</sub>. Thank you for the suggestion.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      We appreciate your suggestion. As suggested by Reviewer #1 also, we will check the phenotype of Polκ knockdown H357-S cells.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      Thank you for the suggestion, we will modify the text accordingly.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      Please note that GFP-Polκ has been overexpressed in H357 Polκ knockout cells to nullify the effect of endogenous Polκ, otherwise we will not be able to test the role of various Polκ mutants.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      We are extremely sorry for the lack of clarity. The inference has been derived from two sets of experiments as shown in Fig. 4C and Fig. 4D (and is with HU).

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      It has already been demonstrated in Fig. 9C where by knocking down USP18, the DDR proteins like Chk1, Chk2, CtIP, and Artemis can be recovered for ubiquitin-mediated proteasomal degradation. The same results are also obtained when its interacting partner Polκ is deleted. In our view, the presented results have sufficiently demonstrated the role of Polκ-Usp18 in the repair of cisplatin adducts through DDR proteins.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

      We appreciate your suggestion. Since the Usp18 protein is not readily available, we will not be able to show; however, we believe the interaction is direct, and we will be able to map the binding site in Polκ.

    1. eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.