10,000 Matching Annotations
  1. Last 7 days
    1. eLife Assessment

      This manuscript reports an important new statistical method for calculating the significance of correlations between two time-series, which provides more accuracy than other methods when the data has few replicates. The proposed method solves a real-life problem that is frequently encountered and is broadly applicable to many realistic datasets in many experimental contexts. The technique is supported with compelling mathematical derivations as well as analysis of both computer-generated and previously published experimental data.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review: Definitions and terminology have been made more precise. Additional analysis confirms the conclusions previously stated and clarifies concerns about the computational tractability of the method.]

      Summary:

      The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore, a useful and broadly applicable statistical tool is well placed in this journal.

      Strengths:

      The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

      The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

      The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

      The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

    3. Reviewer #2 (Public review):

      Summary:

      This paper presented a hypothesis testing procedure for the independence of two time-series that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

      Strengths:

      The test offers good flexibility for different kinds of dependence (through adjusting \rho) and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used, and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore a useful and broadly applicable statistical tool is well placed in this journal.

      Strengths:

      The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

      The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

      The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

      The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

      We appreciate the reviewer’s summary of our study and its strengths.

      Weaknesses:

      A few changes could be made if the manuscript is revised. I would consider all of these points minor, but the paper could be improved if these points were addressed.

      (1) The caption of Figure 2 doesn't seem to mention panel D. Figure A-2 also does not mention C in the caption.

      We apologize for this error, and thank you for catching it! The figure legends had missing or incorrect panel labels. This error has been corrected.

      (2) Figure 2D is a little hard to follow. First, the definition of "Power" is not clear, and I couldn't find the precise definition in the text. Second, the legend for the different lines in 2D is only given in Figure A-2. Perhaps a portion of the caption for Figure 2 is missing?

      We have added a definition of power in the main text:

      “Although the permutation test, simultaneous permute-match test, and sequential permute-match test are all valid, they vary in power – the probability of detecting true dependence.”

      We have clarified the use of “power” in legend of Fig 2 and clarified that the color key for Fig 2D is in Fig 2A. The relevant excerpt of the Fig 2 legend is copied here:

      “(D) Statistical power for the permutation test and various permute-match tests as a function of the replicate number n, significance level α, and strength of dependence r<sub>X, Y</sub>. Power was estimated as the proportion of simulations in which dependence was detected, calculated from 5000 simulations at each value of r<sub>X, Y</sub> between r<sub>X, Y</sub> = 0 and 0.54 in steps of size 0.01. At r<sub>X, Y</sub> = 0, there is no dependence, so the curve at that point indicates the false positive rate rather than power. We chose the Pearson correlation coefficient as our correlation function ρ. See (A) for the color legend.”

      We have also added dotted lines connecting the legend in panel A to the curves in panel D.

      (3) The concept of circular variance for the fish data was heard to understand/visualize. The equation on line 326 did not help much. If there is a very simple picture that could be added near line 326 that helps to explain Ct and theta, that could be a big help for some readers who do not work on related systems. The analysis performed is understandable, the reader just has to accept that circular variance captions the degree of alignment of the fish.

      We have replaced references to circular concentration with “mean resultant length”, which is the standard jargon for this term in circular statistics, and we have added an illustration.

      (4) For the data discussed in Figure 3, I wasn’t 100% sure how the time windows were selected. In the caption, it says “time series to different lengths starting from the first frame”. So the 20 s time window was from t=0 to t= 20 s. Would a different result be obtained if a different 20 s window was chosen (from t = 4 min to t = 4 min 20 s just to give a specific example). I suppose by chance one of the time windows would give a pvalue less than the target 0.05, that wouldn’t be surprising. Maybe a random time window should be selected (although I am not indicating what was reported was incorrect)? A little more discussion on this aspect of the study may be helpful.

      As suggested by the reviewer, we have redone the analysis of Figure 3D with random segments. This provides a more complete picture of how the chance of detecting a significant correlation varies with segment length. The main conclusion is unchanged: Perfect match tests reliably detect dependence across a wider range of segment lengths than the naive parametric alternative.

      The relevant panel and an excerpt from the legend text are copied below.

      “(D) Permute-match tests detected a significant correlation between speed and alignment more consistently than the parametric test. For a grid of lengths between 20 and 600 seconds we sampled 500 random segments of each length, each drawn from the first 600 seconds, and determined for each segment whether the parametric test and/or the two possible permute-match tests detected a significant (p ≤ 0.05) correlation. In the edge case of the maximum 600-second length, all 500 “random” segments were identical.”

      Reviewer #2 (Public review):

      Summary:

      This paper presented a hypothesis testing procedure for the independence of two timeseries that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

      Strengths:

      The test offers good flexibility for different kinds of dependence (through adjusting \rho), and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

      We appreciate the reviewer’s summary of key aspects of our manuscript.

      Weaknesses:

      (1) The size of the test is not guaranteed to (asymptotically) equal \alpha, which may damage the power.

      We thank the reviewer for raising the issue of test size and power. We agree that a conservative test (one whose size can fall below alpha) may sacrifice power.

      Our objective is distribution-free false-positive rate (FPR) control. That is, we wish to keep the FPR at or below alpha for every distribution of X and Y, because in our regime (nonstationary time series with few independent replicates) the scientist often cannot verify distributional assumptions. Inspired by the reviewer’s comment, we now show (new Proposition 14) that the perfect match probability can be made arbitrarily close to 1/n<sup>!</sup>. As a consequence, any reported perfect match p-value below 1/n<sup>!</sup> would break the distribution-free validity of the test.

      A test that exploits distributional structure could likely access lower p-values; we have now explored how the empirical FPR of the permute-match test varies with the data-generating process (see our response to reviewer 2's recommendation 1 below).

      (2) The computational time can be an issue for a moderately large sample size when calculating the X / Y-perfect match. It will be beneficial to include discussions on the implementations of the test.

      We agree this is an important consideration. We have added the following text to the Discussion:

      “The test appears computationally tractable for relevant sample sizes: Our implementation of the permute match procedure completed a single test of dependence in the setting of Fig 2 with an average runtime of 3 seconds when n = 10 on a 2023 14-inch MacBook Pro with an M2 Pro processor and 16 GB RAM (see Source data 1). For n > 10, a standard permutation test already can report a p-value below 3 × 10<sup>−8</sup> so the perfect match test is likely unnecessary for typical applications.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A few more minor notes/comments:

      As a personal preference, I like it when figures printed in grayscale retain their meaning (when possible). Just FYI, Figure 3D in grayscale is uninterpretable. Not saying a change is needed, just pointing it out.

      We changed Fig 3D to address a comment above, and think the new version better distinguishes between the permute-match and permutation test results in greyscale.

      On line 304, is that a lower bound or an upper bound? Maybe the issue is the probability mentioned on line 305 is not clear.

      As this point is not the main focus of the investigation, we have rephrased it to make it less technical and eliminate the issue of which bound is in question.

      “It seems likely that the tests could be further modified to report an even lower p-value when an X- and Y -perfect match occur simultaneously, as in Fig 3B. However, we have not investigated further and this problem is left for future efforts.”

      I am not sure if this is a weakness, but the p-value changing depending on the choice of whether to apply the X-perfect match or Y-perfect match test first is fascinating. The authors did discuss this very issue at several points in the manuscript. It is slightly unsettling to me that there isn't an exact p-value for a given set of data. This one point gives me a new perspective on statistics.

      In full transparency, I don't believe I have to background to thoroughly review the appendix. I did read through it and did not notice any errors, but I couldn't confidently say there are not any small mathematical errors or any logical flaws in the proofs. Some sections were not easy to follow (my own shortcomings, the writing appeared sufficient for more of an expert to understand).

      We greatly appreciate the reviewer’s time and effort tackling an appendix outside their comfort zone.

      Reviewer #2 (Recommendations for the authors):

      (1) In the numerical experiment session, the authors should include the null situation, i.e., the performance of the test when X and Y are independent. This helps assess the size of the test.

      We have added a section on size to our results section, copied below:

      “The permute-match test’s false positive rate depends on the process tested. The permute-match test is conservative – meaning that its false positive rate can fall below the significance level – because both the permutation test and perfect match test are conservative. As discussed elsewhere [29], the permutation test is conservative when α is not one of its possible p-values and when ties may occur between the original correlation and shuffled correlations. Checking for a perfect match is similarly conservative. The actual probability of a false-alarm perfect match event can vary depending on the process being tested. To see this consider the permute-match test in the setting where n = 3, where α = 0.05, and where r<sub>X,Y</sub>= 0 (independent X and Y). Note that in this case, obtaining a Y -perfect match (and thus p = 1/n<sup>n</sup>) is necessary and sufficient to detect dependence since α is too low for detection by either the permutation test or the p = 2/n<sup>n</sup> leg of the permute-match test. In the linear system of Fig 2, we observed among 5000 simulations a detection rate of 0.0148, significantly below the upper bound of 1/3<sup>3</sup> (Figure 2 - Source data 1; one-tailed exact binomial test, p < 10<sup>−20</sup>). Conversely, in the nonlinear system of Fig S2, this same event (Y -perfect match under n = 3 and r<sub>X,Y</sub> = 0) occurs with a detection rate of 0.0328, not significantly 9 below the upper bound of 1/3<sup>3</sup> (Figure S2 - Source data 1; one-tailed exact binomial test, p = 0.059). Thus, depending on the underlying process studied, the actual chance of a perfect match happening under independence may be near or significantly below the theoretical upper bound.”

      (2) Some insights regarding the choice of rho should be provided. Especially, are there any examples that the classical test, such as the Pearson correlation or Granger causality test does not work?

      We have redone the example of Appendix 4 with Pearson correlation, showing that Pearson correlation has substantially lower power than cross-map skill in this case (compare figures S2 and S3).

      (3) Line 34 - 35, page 2: Correlations and causality should be separately considered. This sentence talks more about causality rather than correlation.

      We appreciate the reviewer’s perspective and agree that correlation and causality are distinct.

      We feel that pointing out the issue of spurious correlations is helpful to orient our readers, especially those from a broad scientific audience. In the text, we define “correlation” as a descriptive statistic (rather than normalized covariance), and later distinguish it from “dependence”, which has causal implications due to Reichenbach’s common cause principle. We believe this distinction provides a useful backdrop for practitioners who use statistical methods but are perhaps new to thinking deeply about statistical dependence.

      (4) Please add some discussions on the situation that X_i depends on Y_{i - j} for some j > 0, which is associated with the setting of Granger causality test.

      We have added the following to the discussion:

      “No distributional assumptions are required, and the correlation function ρ can be completely arbitrary. For instance, ρ could include a lag to detect delayed dependence, or even evaluate the correlation strength at several lags and report the strongest among them [39].”

    1. eLife Assessment

      This important study by Otgonbaatar and colleagues employs advanced live microscopy, optogenetics, and an endogenous fluorescent timer system to investigate short- and long-term stabilization dynamics of β-catenin/Armadillo (Arm) during Drosophila development. The authors identify an unexpected and functionally relevant enrichment of stabilized junctional Arm in leading-edge cells during dorsal closure, providing evidence for a stabilization mechanism that appears independent of canonical Wingless signaling. These findings are significant because they expand current understanding of β-catenin/Arm beyond its canonical signaling functions and suggest a role in tissue mechanics and force transmission during dorsal closure. The proposed model represents a key advance in the field, but the strength of evidence is currently incomplete: the main conclusions regarding Wingless independence, JNK-mediated regulation, and the mechanical role of stabilized Arm are only partially supported by the available data and would benefit from further experimental testing and corroboration.

    2. Reviewer #1 (Public review):

      In this study, Otgonbaatar and colleagues investigate the stability of Armadillo (Arm) during Drosophila development using a creative tandem fluorescent protein timer approach via endogenous tagging of Arm. The tagging strategy allows for newly synthesised and longer-term stabilised Arm pools to be distinguished from one another. Specifically, the authors address the functional relevance of and mechanism behind the stabilisation of junctional Arm during dorsal closure.

      The authors show that Arm is stabilised at the leading edge during dorsal closure. Using a sophisticated optogenetics approach, which allows for acute perturbations, they show that stabilised Arm is functionally required for dorsal closure. Increasing Wg (by overexpression) did not affect dorsal closure or Arm stability, in contrast to Axin overexpression, which reduces Wg/Arm signalling. In line with canonical signalling control of Arm levels being critical, stabilisation of Arm by N-terminal mutations disrupted dorsal closure. However, the same deletion is also expected to affect interaction with alpha-catenin. Co-localisation with E-cadherin and actin suggests a junctional role of leading-edge localised Arm. Optogenetic targeting of alpha-catenin points towards a key role of adherence junctions in dorsal closure. Allele replacement with mutant variants of Arm to affect adherence junction complex assembly further indicates an important contribution of coupling between Arm and alpha-catenin. Using overexpression approaches, the authors suggest that Dsh and Jnk contribute to dorsal closure.

      This microscopy- and optogenetics-based study is generally well-conducted and provides strong evidence for stabilised Arm during dorsal closure, as well as its functional importance. This is an important discovery relevant to morphogenesis and potentially mechanotransduction. From a technical perspective, the validated beta-catenin timer provides a valuable tool for the field. The timer has revealed that Arm stabilisation does not coincide with Wg stripes, suggesting a Wg-independent stabilisation mechanism that may instead depend on adherence junction assembly, especially the interaction of Arm with alpha-catenin. However, as N-terminal deletion within Arm and Axin overexpression also disrupted dorsal closure, substantial ambiguity remains. Can suppression of the beta-catenin degradation machinery be ruled out as a regulatory mechanism? An expansion of ArmTimer mutant variants could contribute to testing the authors' conclusion further. Structural insights into junctional interactions involving Arm (e.g., 10.1074/jbc.M114.554709) could, for example, be used for further functional exploration by mutagenesis. The direct mechanistic impact of JNK and its potential link to Dsh in dorsal closure remains less compelling.

      In summary, this is a highly relevant and important study, potentially pointing to a novel stabilisation mechanism of beta-catenin in development. Further corroboration of the mechanism, to test whether it is indeed distinct from canonical signalling, would be needed to support the conclusions.

    3. Reviewer #2 (Public review):

      Summary:

      Otgonbaatar et al. sought to investigate β-catenin/Arm protein lifetime and stabilization dynamics in vivo during embryonic development. To address this question, the authors developed an endogenous tandem fluorescent protein timer (tFP) system that enables the visualization of newly synthesized versus long-lived Arm protein in vivo. Using this approach, the authors sought to determine where stabilized Arm accumulates during development and how it contributes to dorsal closure.

      Strengths:

      A major strength of the study is the development and application of the endogenous Arm timer system, which provides a powerful approach for monitoring protein stabilization dynamics in living tissues. Using this system, the authors unexpectedly found that the strongest Arm stabilization occurs not in Wnt signaling regions, but at the leading edge cells during dorsal closure. The study combines quantitative live imaging, optogenetic perturbation, genetic analysis, and structure-function approaches to demonstrate that stabilized junctional Arm interacts with α-catenin and contributes to tissue mechanics required at the leading edge for dorsal closure. Particularly compelling is the combination of multiple perturbations, including optogenetic disruption of Arm or α-catenin, Axin overexpression, and Arm mutants, which produce consistent dorsal closure defects.

      Some conclusions are generally supported by the presented data. The work provides strong evidence that Arm plays an important role in dorsal closure. The identification of a requirement for the Dishevelled DEP domain and JNK signaling supports a non-canonical regulatory mechanism controlling dorsal closure.

      Weaknesses:

      (1) Conclusions are made regarding force transmission;(however, no experimental evidence is provided to support these conclusions.

      (2) The conclusion was made that Wingless does not affect dorsal closure. However, this was based solely on Wingless overexpression in the amnioserosa, and the level of Wingless expression was not quantified. One possibility is that this level was not sufficient to see an effect. Alternatively, Wingless may have a role in migrating epithelium rather than the amnioserosa. Indeed, it is known that wingless mutants display a defect in dorsal closure.

      (3) The effect of JNK knockdown on Arm localization maybe is indirect, and due to a secondary consequence on disruption of epithelial morphology rather than a direct effect of JNK on Arm.

      (4) Some conclusions rely on overexpression-based perturbations (e.g., Axin or Arm mutants), which may not fully recapitulate endogenous physiological regulation.

      (5) The Arm timer was not able to detect Wingless-dependent Arm stabilization in stripes. This finding demonstrates that the timer is not sensitive enough to thoroughly analyze Arm dynamics.

      Overall, this work provides important conceptual advances in understanding junctional β-catenin/Arm function during dorsal closure. The endogenous fluorescent timer approach will likely be broadly useful to the community for studying protein stability dynamics in vivo, and the findings expand current views of β-catenin by highlighting its mechanical and junctional functions during tissue morphogenesis.

    4. Author response:

      We are glad the reviewers found the tandem fluorescent timer approach valuable and the leading-edge Arm stabilization finding significant.

      We agree with the Assessment that our evidence for three specific claims Wingless-independence, JNK-mediated regulation of Arm stability, and a direct mechanical/force-transmission role for stabilized Arm is currently incomplete, and we will revise the text throughout to reflect this more precisely rather than overstating the current data. In addition, we commit to two new experiments, both using existing reagents and fly stocks, that speak directly to the two most experimentally tractable points raised by the reviewers:

      (1) Re-staining our existing JNK-RNAi and JNK-overexpression embryos for E-cadherin (reagent already validated in Figure 4), to test whether JNK acts directly on junctional architecture or only indirectly, via broader epithelial disruption.

      (2) Imaging ArmTimer in a wingless loss-of-function background, to directly test Wingless-dependence of leading-edge Arm stabilization as the reciprocal of our existing overexpression data.

      We address each public review point below and outline the accompanying text revisions.

      On Wingless independence (Reviewer #1; Reviewer #2, Weaknesses #2 and #5; Recommendation #1):

      We agree that our current evidence unquantified Wg overexpression restricted to the amnioserosa (C381-Gal4) and uniform overexpression, alongside the absence of detectable Wg-stripe-associated Arm-Timer signal supports a more limited conclusion than "Wingless-independent" as currently stated. We will revise our language throughout the Abstract, Results, and Discussion to state that canonical Wg overexpression does not detectably enhance leading-edge Arm stabilization or perturb dorsal closure under our conditions, rather than asserting pathway independence. As noted above, we commit to imaging ArmTimer in a wg mutant background to test this directly, complementing our overexpression data with the reciprocal loss-of-function manipulation.

      On the related point that the Timer's failure to detect a Wg-stripe-associated stabilization signal could reflect a sensitivity limitation rather than a true absence of stabilization (Reviewer #2, Weakness #5): we agree and will state this explicitly rather than treating absence of signal as evidence of absence. This does not undermine the positive leading-edge finding, which is not defined relative to the stripe comparison: all embryos, channels, and time points were imaged and rendered using identical laser power and brightness/sensitivity settings, and the leading-edge RFP signal clearly exceeds background under those same acquisition conditions. We also note that detection limits of this kind are a recognized challenge for endogenously tagged reporters of canonical Wnt/β-catenin signaling generally, including in mammalian systems, and cite two studies already in our bibliography that report the same class of limitation: de Man et al. (2021, eLife 10:e66440) and Ambrosi et al. (2022, eLife 11:e64498). We have added this clarification, with these citations, to the Results (paragraph describing Figure 3).

      On the mechanistic link between Arm stability and destruction-complex activity (Reviewer #1):

      We agree that because both ΔArm and Axin overexpression converge on the destruction complex, our data cannot yet fully separate "escape from degradation" from "impaired α-catenin/junctional coupling" as the operative mechanism. We will revise the Discussion to state this ambiguity explicitly and will treat the ArmTimer-AA result (partial α-catenin-binding disruption via phosphosite mutation, independent of destruction-complex regulation) as the strongest current evidence isolating the junctional-coupling mechanism. We also thank Reviewer 1 for pointing us to Pokutta, Choi, Ahlsen, Hansen & Weis (2014, J Biol Chem 289:13589-13601), which structurally and thermodynamically characterized the mammalian cadherin·β-catenin·α-catenin complex and showed that α-catenin binding to β-catenin is a distinct, allosterically regulated interface cadherin binding increases β-catenin's affinity for α-catenin roughly 10-fold, and α-catenin homodimerization independently competes with β-catenin binding. We have added this citation to the Discussion as structural support for treating cadherin engagement, α-catenin coupling, and destruction-complex regulation as mechanistically separable interfaces, and note that the crystallized β-catenin·α-catenin interface provides a structural template for future experiments for example, structure-guided point mutations at the homologous interface residues in Arm, or in vitro binding assays comparing wild-type and threonine-mutant (T111A/T121A) Arm affinity for α-catenin.

      We do not, however, believe a destruction-complex-independent stabilizing allele of Arm is a tractable experiment to close this gap directly: any allele that stabilizes Arm without engaging the destruction complex is, by definition, a Wnt pathway gain-of-function allele, since destruction-complex-mediated degradation is the very regulatory step that canonical Wnt signaling controls. Nor would restricting the allele to a transcriptionally inactive form of Arm cleanly resolve the confound: Wnt/TCF target loci include dedicated repressive TCF-binding sites (Blauwkamp, Chang & Cadigan, 2008, EMBO J 27:1436-1446), so a transcriptionally "dead" stabilized Arm could still alter transcription by disrupting TCF-mediated repression. We therefore treat this as a genuine, currently unresolvable confound of the overexpression approach, and rely instead on the CRY2 optogenetic and ArmTimer-AA results as the strongest available evidence isolating a junctional-coupling contribution. We have added this reasoning, with both citations, to the Discussion.

      On JNK acting on Arm directly vs. indirectly (Reviewer #1; Reviewer #2, Weakness #3 and Recommendation #2):

      This is the most actionable point raised by both reviewers. As noted above, we commit to re-imaging and re-staining our existing JNK-RNAi and JNK-overexpression embryos for E-cadherin to determine whether junctional/polarity architecture is broadly disrupted under these conditions (indirect mechanism) or whether E-cadherin localization is comparatively preserved while Arm stabilization is specifically altered (direct mechanism). In the meantime, we note that a direct mechanism is biochemically plausible: in mammalian cells, JNK phosphorylates β-catenin directly and regulates adherens junction integrity, and JNK activity separately controls the binding of α-catenin to the junctional complex (Lee, Koria, Qu & Andreadis, 2009, FASEB J 23:3874-3883; Lee, Padmashali, Koria & Andreadis, 2011, FASEB J 25:613-623). We cite these as precedent that a direct route from JNK to junctional β-catenin/α-catenin regulation exists in another system, while being explicit that this does not establish the same mechanism in Drosophila dorsal closure that will be tested directly by the E-cadherin re-staining experiment. We have added these citations and this caveat to the Discussion.

      On the Dsh-DEP-to-JNK mechanistic link (Reviewer #1, Public Review #2 and Recommendation #5):

      We agree that our data show the Dsh-DEP requirement and the JNK requirement for dorsal closure as parallel, independent findings rather than a demonstrated linear pathway in our system. To provide context for why we consider a DEP-to-JNK connection a reasonable working hypothesis, we searched the literature in both Drosophila and vertebrates and will cite six additional studies establishing this link: Axelrod et al. (1998) and Axelrod (2001), establishing that DEP-dependent membrane recruitment and unipolar localization of Dishevelled are specifically required for planar polarity signaling, distinct from Wingless signaling; Paricio et al. (1999) and Fanto et al. (2000), showing Dishevelled acts through Misshapen and Rac1/RhoA to the same JNK module used in dorsal closure; and Moriguchi et al. (1999) and Yamanaka et al. (2002), showing biochemically in vertebrates that the DEP domain of Dvl-1 selectively activates JNK independent of β-catenin/TCF-LEF activity, and that this JNK requirement is conserved in Xenopus convergent extension, the vertebrate process most functionally analogous to dorsal closure. We will state explicitly that this precedent, while now cross-species, comes from planar-cell-polarity and convergent-extension assays rather than dorsal closure itself, so it supports the plausibility of a Dsh/Dvl-DEP-to-JNK connection without establishing that the identical pathway operates in our system.

      On the Dsh DIX/DEP domain-separability argument (Reviewer #1, Recommendation #4):

      We thank the reviewer for pointing us to Gammons, Renko, Johnson, Rutherford & Bienz (2016, Mol Cell 64:92-104), which showed that the Wnt signalosome itself is assembled by head-to-tail DEP domain swapping between Dishevelled molecules, and that this DEP-dependent oligomerization is directly required for canonical Wnt pathway activity not restricted to the non-canonical/planar-polarity branch as we had implied. We agree this evidence undercuts our previous interpretation of the DshΔDEP dorsal closure phenotype as evidence for a strong non-canonical/polarity-specific role for the DEP-dependent branch of Dsh. We have revised the Discussion accordingly: we now state that the DEP domain is required for the morphogenetic program culminating in dorsal closure, cite Gammons et al. directly, and note that this requirement does not by itself establish a non-canonical/polarity-specific role, since we cannot rule out a contribution from DEP-dependent canonical Wnt signalosome assembly.

      On force transmission (Reviewer #2, Weakness #1):

      We agree that we have not directly measured force or tension at the leading edge, and that our current data (colocalization with actin/E-cadherin, and functional requirement shown via CRY2 optogenetics and mutant analysis) are consistent with, but do not directly demonstrate, a role in force transmission. We do not have the in-house expertise to perform direct force/tension measurements (e.g., laser ablation, junctional tension assays), so we will not be adding such an experiment in this revision. Instead, we have revised the language throughout the manuscript including two Discussion section headings that previously stated a mechanical role for stabilized Arm as established fact to consistently present the mechanical/force-transmission role as a hypothesis raised by our data, not a demonstrated conclusion, and we retain a clear statement that direct force measurement (ideally in collaboration with groups with the relevant biophysical expertise) is future work rather than a claim we are making in this manuscript.

      On the phosphomimetic threonine mutant (Reviewer #1, Recommendation #2):

      ArmTimer-AA (T111A, T121A) was generated with the expectation that the tyrosine phosphosite mutants (ArmTimer-EE, ArmTimer-FF) would be the primary drivers of any dorsal closure phenotype, given their proposed role in E-cadherin binding; the pronounced zippering defect we observed in ArmTimer-AA was therefore an unanticipated finding rather than a predicted result. We have not generated the reciprocal phosphomimetic ArmTimer-EE(Thr) (T111E, T121E) allele. Generating and characterizing this allele is a substantial undertaking we estimate over a year including allele generation, validation, and phenotypic characterization and we will state this explicitly in the Discussion as planned future work rather than part of the current revision.

      On confirmation of myristoylated-Dsh membrane targeting (Reviewer #1, Recommendation #3):

      We cannot confirm that myristoylation localizes all Dsh protein to the membrane. However, this strategy has extensive prior genetic validation using the identical Src-derived myristoylation sequence: it was originally used to tether Armadillo and shown sufficient for constitutive Wnt pathway activation (Zecca, Basler & Struhl, 1996; Tolwinski & Wieschaus, 2001, 2004), and the same approach was subsequently applied to GSK3 and Dishevelled, in each case producing the expected pathway-activation phenotypes (Mannava & Tolwinski, 2015; Kaur et al., 2017). We have added these citations to the Results where the Myr-Dsh constructs are introduced.

      On overexpression-based perturbations versus endogenous regulation (Reviewer #2, Weakness #4):

      We would like to clarify that most of the Arm alleles used in this study including all of the point-mutant Timer alleles (ArmF1a, ArmTimer-FF, ArmTimer-EE, ArmTimer-AA) central to our mechanistic conclusions were generated as knock-ins at the endogenous ‘arm’ locus via MiMIC/RMCE, not overexpressed. The two exceptions are ΔArm and ArmS56A, expressed from UAS constructs because both are gain-of-function alleles anticipated to be lethal if expressed from the endogenous locus, based on prior experience with similarly stabilizing mutations. Axin overexpression was used because no Axin mutant or knock-in allele was generated for this study; we agree an endogenous Axin allele would be the ideal complement and will state this explicitly as a limitation, while noting that our CRY2 optogenetic perturbations of Arm and α-catenin which act acutely on the endogenous proteins provide an orthogonal line of evidence supporting the same conclusions. We have added this clarification to the Discussion.

    1. eLife Assessment

      This study presents a valuable RNA velocity solution which integrates cell differentiation and gene regulation, with a balance between neuralODE and raw gene space. The evidence supporting the claims of the authors is solid, although inclusion of discussion on the challenges in capturing cell cycle transitions would have strengthened the study. The work will be of interest to scientists working in the field of computational biology and gene regulation.

    2. Reviewer #1 (Public review):

      Summary:

      In the paper, the authors propose a new RNA velocity method, TSvelo, which predicts the transcription rate linearly based on the expression of RNA levels of transcription factors. This framework is an extension of its recent work TFvelo by including unspliced reads and designing a coherent neuralODE framework. Improved performance was demonstrated in six diverse datasets.

      Strengths:

      Overall, this method introduces innovative solutions to link cell differentiation and gene regulation, with a balance between model complexity (neuralODE) and interpretability (raw gene space).

      Comments on revised version:

      I thank the authors for further revision, and I do not have any other concerns. I believe it is an important contribution to this field of trajectory inference and gene regulation.

    3. Reviewer #3 (Public review):

      Despite the abundance of RNA velocity tools, there are still major limitations, and there is strong skepticism about the results these methods lead to. In this paper, the authors try to address some limitations of current RNA velocity approaches by proposing a unified framework to jointly infer transcriptional and splicing dynamics. The method is then benchmarked on 6 real datasets against the most popular RNA velocity tools.

      Comments on revised version:

      The Authors addressed my 2 follow-up comments suitably.

      Thanks for the time you took addressing them. I have no further comments.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the paper, the authors propose a new RNA velocity method, TSvelo, which predicts the transcription rate linearly based on the expression of RNA levels of transcription factors. This framework is an extension of its recent work TFvelo by including unspliced reads and designing a coherent neuralODE framework. Improved performance was demonstrated in six diverse datasets.

      Strengths:

      Overall, this method introduces innovative solutions to link cell differentiation and gene regulation, with a balance between model complexity (neuralODE) and interpretability (raw gene space).

      Comments on revised version:

      The authors have added comprehensive analyses in this revision, and all of my concerns have been very well addressed. Here, I just want to re-emphasize the original points 1 and 3.

      (1) The analysis and clarification are very helpful - thanks! I found that Fig. R1 and R2 are very insightful, as DoRothEA-only returns much worse performance. Please consider adding these two figures to the supp figure and possibly highlighting your setting for edge pruning (down-weights); therefore, the model is more likely to be affected by false negatives than false positives in the TF-target prior.

      We thank the reviewer for the positive feedback and for recognizing the value of the additional analyses. We have added the previous Fig. R1 and Fig. R2 to the Supplementary Information as Fig. S13 and Fig. S14, respectively, and have referred to them in the revised manuscript.

      We have also expanded the description of the TF–target prior used in TSvelo in the “Acquiring Prior Knowledge of Gene Regulatory Relations” subsection of the Methods. As noted by the reviewer, TSvelo is expected to be less sensitive to false-positive TF–target interactions because unsupported edges can be down-weighted during training. In contrast, missing true regulatory interactions are not represented in the prior network and therefore cannot contribute to the learned regulatory dynamics, making the model potentially more sensitive to false negatives.

      (3) Please consider adding some discussion on the challenges in capturing cell cycle transitions.

      We thank the reviewer for this suggestion. We have added a brief discussion in the Discussion section on the challenges of modeling cell-cycle transitions. In particular, cell-cycle progression is often characterized by cyclic dynamics and overlapping transcriptional programs, which can complicate the inference of directional state transitions and regulatory relationships.

      Reviewer #3 (Public review):

      Despite the abundance of RNA velocity tools, there are still major limitations, and there is strong skepticism about the results these methods lead to. In this paper, the authors try to address some limitations of current RNA velocity approaches by proposing a unified framework to jointly infer transcriptional and splicing dynamics. The method is then benchmarked on 6 real datasets against the most popular RNA velocity tools.

      Comments on revised version.

      The Authors addressed all my comments suitably. I'd like to thank them for the time they spent addressing them: the revised paper is much more convincing.

      I have 2 very minor follow-up concerns:

      (1) I appreciated the simulation study, however, no null simulation is present.

      We know RNA velocity tools are inclined to provide false positives: trajectories even when the data doesn't have any.

      I'd be helpful to add null simulations where the data has no trajectories and see if methods erroneously identify any.

      We thank the reviewer for this helpful suggestion. We have added null simulations to evaluate TSvelo and baseline approaches on data without underlying dynamic structure. Specifically, we generated a null dataset including 200 genes and 600 cells by independently sampling spliced (S) and unspliced (U) counts, thereby removing any coherent transcriptional relationship between them.

      When applying scVelo and UniTVelo to this data, no genes passed the velocity gene selection step under the default likelihood-based filtering, and no velocity field could be obtained. We further tested TSvelo, Dynamo, and cellDancer on the same null data and observed that all three methods still produce trajectory-like patterns despite the absence of true dynamics (See Supplementary Information as Fig. S17).

      Including TSvelo, many RNA velocity and trajectory inference approaches assume that they are applied to datasets reflecting underlying dynamic biological processes. We agree that incorporating additional checks during preprocessing could help prevent applying velocity analysis to non-dynamic datasets. We have added this discussion to the revised manuscript.

      (2) Several of the novel analyses are only reported in the Supplementary material and only references in the main text (e.g., "A validation of TSvelo on simulated data is provided in Fig. S1 and Fig. S2 in the Supplementary Information."). This is pity!

      If allowed, I'd add some comments about the new analyses (simulations, computational benchmarks, etc...) also in the main text.

      We thank the reviewer for this suggestion. We agree that several analyses presented in the Supplementary Information provide important support for our conclusions. To improve their visibility, we have expanded the corresponding descriptions in the main text and briefly summarized the key findings of the relevant Supplementary Figures instead of only citing them. These revisions have been made for Fig. S1, Fig. S2, Fig. S10, Fig. S12, Fig. S13, Fig. S14 and Fig. S17. In particular, we have incorporated a summary of the simulation results at the end of the subsection “Estimate RNA Velocity with TSvelo” in the Results section, and added a discussion of the computational benchmarking analyses in the Discussion section. We hope these changes improve the accessibility of these results while maintaining a concise presentation of the main findings.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      I suggest the paper to undergo (very) minor revisions as detailed in the Public Review.

      Simone Tiberi, The University of Bologna

      We sincerely thank all reviewers for their thoughtful suggestions, which have helped improve the clarity and overall presentation of the manuscript.

    1. eLife assessment

      In this manuscript, Rademacher and colleagues examined the effect of a chemogenetic approach on the integrity of the dopamine system in mice with chronically stimulating dopamine neurons. These findings are important: (1) This approach led to an axon-first degeneration over a time course (2–4 weeks) that is suitable for experimental investigation; (2) The finding that direct excitation of dopaminergic neurons causes differential degeneration sheds light on dopaminergic neuron selective vulnerability mechanisms. Overall, the strength of the evidence is solid, but the behavior experiments that do not include a CNO control provide incomplete support for the findings.

    2. Reviewer #1 (Public Review):

      Summary:

      In this manuscript, the authors investigated the effect of chronic activation of dopamine neurons using chemogenetics. Using Gq-DREADDs, the authors chronically activated midbrain dopamine neurons and observed that these neurons, particularly their axons, exhibit increased vulnerability and degeneration, resembling the pathological symptoms of Parkinson's disease. Baseline calcium levels in midbrain dopamine neurons were also significantly elevated following the chronic activation. Lastly, to identify cellular and circuit-level changes in response to dopaminergic neuronal degeneration caused by chronic activation, the authors employed spatial genomics (Visium) and revealed comprehensive changes in gene expression in the mouse model subjected to chronic activation. In conclusion, this study presents novel data on the consequences of chronic hyperactivation of midbrain dopamine neurons.

      Strengths:

      This study provides direct evidence that the chronic activation of dopamine neurons is toxic and gives rise to neurodegeneration. In addition, the authors achieved the chronic activation of dopamine neurons using water application of clozapine-N-oxide (CNO), a method not commonly employed by researchers. This approach may offer new insights into pathophysiological alterations of dopamine neurons in Parkinson's disease. The authors also utilized state-of-the-art spatial gene expression analysis, which can provide valuable information for other researchers studying dopamine neurons. Although the authors did not elucidate the mechanisms underlying dopaminergic neuronal and axonal death, they presented a substantial number of intriguing ideas in their discussion, which are worth further investigation.

      Weaknesses:

      Many claims raised in this paper are only partially supported by the experimental results. So, additional data are necessary to strengthen the claims. The effects of chronic activation of dopamine neurons are intriguing; however, this paper does not go beyond reporting phenomena. It lacks a comprehensive explanation for the degeneration of dopamine neurons and their axons. While the authors proposed possible mechanisms for the degeneration in their discussion, such as differentially expressed genes, these remain experimentally unexplored.

    3. Reviewer #2 (Public Review):<br /> <br /> Summary:

      Rademacher et al. present a paper showing that chronic chemogenetic excitation of dopaminergic neurons in the mouse midbrain results in differential degeneration of axons and somas across distinct regions (SNc vs VTA). These findings are important. This mouse model also has the advantage of showing a axon-first degeneration over an experimentally-useful time course (2-4 weeks). 2. The findings that direct excitation of dopaminergic neurons causes differential degeneration sheds light on the mechanisms of dopaminergic neuron selective vulnerability. The evidence that activation of dopaminergic neurons causes degeneration and alters mRNA expression is convincing, as the authors use both vehicle and CNO control groups, but the evidence that chronic dopaminergic activation alters circadian rhythm and motor behavior is incomplete as the authors did not run a CNO-control condition in these experiments.

      Strengths:<br /> This is an exciting and important paper.<br /> The paper compares mouse transcriptomics with human patient data.<br /> It shows that selective degeneration can occur across the midbrain dopaminergic neurons even in the absence of a genetic, prion, or toxin neurodegeneration mechanism.

      Weaknesses:

      Major concerns:

      (1) The lack of a CNO-positive, DREADD-negative control group in the behavioral experiments is the main limitation in interpreting the behavioral data. Without knowing whether CNO on its own has an impact on circadian rhythm or motor activity, the certainty that dopaminergic hyperactivity is causing these effects is lacking.

      (2) One of the most exciting things about this paper is that the SNc degenerates more strongly than the VTA when both regions are, in theory, excited to the same extent. However, it is not perfectly clear that both regions respond to CNO to the same extent. The electrophysiological data showing CNO responsiveness is only conducted in the SNc. If the VTA response is significantly reduced vs the SNc response, then the selectivity of the SNc degeneration could just be because the SNc was more hyperactive than the VTA. Electrophysiology experiments comparing the VTA and SNc response to CNO could support the idea that the SNc has substantial intrinsic vulnerability factors compared to the VTA.

      (3) The mice have access to a running wheel for the circadian rhythm experiments. Running has been shown to alter the dopaminergic system (Bastioli et al., 2022) and so the authors should clarify whether the histology, electrophysiology, fiber photometry, and transcriptomics data are conducted on mice that have been running or sedentary.

    4. Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Rademacher and colleagues examined the effect on the integrity of the dopamine system in mice of chronically stimulating dopamine neurons using a chemogenetic approach. They find that one to two weeks of constant exposure to the chemogenetic activator CNO leads to a decrease in the density of tyrosine hydroxylase staining in striatal brain sections and to a small reduction of the global population of tyrosine hydroxylase positive neurons in the ventral midbrain. They also report alterations in gene expression in both regions using a spatial transcriptomics approach. Globally, the work is well done and valuable and some of the conclusions are interesting. However, the conceptual advance is perhaps a bit limited in the sense that there is extensive previous work in the literature showing that excessive depolarization of multiple types of neurons associated with intracellular calcium elevations promotes neuronal degeneration. The present work adds to this by showing evidence of a similar phenomenon in dopamine neurons. In terms of the mechanisms explaining the neuronal loss observed after 2 to 4 weeks of chemogenetic activation, it would be important to consider that dopamine neurons are known from a lot of previous literature to undergo a decrease in firing through a depolarization-block mechanism when chronically depolarized. Is it possible that such a phenomenon explains much of the results observed in the present study? It would be important to consider this in the manuscript. The relevance to Parkinson's disease (PD) is also not totally clear because there is not a lot of previous solid evidence showing that the firing of dopamine neurons is increased in PD, either in human subjects or in mouse models of the disease. As such, it is not clear if the present work is really modelling something that could happen in PD in humans.

      Comments on the introduction:

      The introduction cites a 1990 paper from the lab of Anthony Grace as support of the fact that DA neurons increase their firing rate in PD models. However, in this 1990 paper, the authors stated that: "With respect to DA cell activity, depletions of up to 96% of striatal DA did not result in substantial alterations in the proportion of DA neurons active, their mean firing rate, or their firing pattern. Increases in these parameters only occurred when striatal DA depletions exceeded 96%." Such results argue that an increase in firing rate is most likely to be a consequence of the almost complete loss of dopamine neurons rather than an initial driver of neuronal loss. The present introduction would thus benefit from being revised to clarify the overriding hypothesis and rationale in relation to PD and better represent the findings of the paper by Hollerman and Grace.

      It would be good that the introduction refers to some of the literature on the links between excessive neuronal activity, calcium, and neurodegeneration. There is a large literature on this and referring to it would help frame the work and its novelty in a broader context.

      Comments on the results section:

      The running wheel results of Figure 1 suggest that the CNO treatment caused a brief increase in running on the first day after which there was a strong decrease during the subsequent days in the active phase. This observation is also in line with the appearance of a depolarization block.

      The authors examined many basic electrophysiological parameters of recorded dopamine neurons in acute brain slices. However, it is surprising that they did not report the resting membrane potential, or the input resistance. It would be important that this be added because these two parameters provide key information on the basal excitability of the recorded neurons. They would also allow us to obtain insight into the possibility that the neurons are chronically depolarized and thus in depolarization block.

      It is great that the authors quantified not only TH levels but also the levels of mCherry, co-expressed with the chemogenetic receptor. This could in principle help to distinguish between TH downregulation and true loss of dopamine neuron cell bodies. However, the approach used here has a major caveat in that the number of mCherry-positive dopamine neurons depends on the proportion of dopamine neurons that were infected and expressed the DREADD and this could very well vary between different mice. It is very unlikely that the virus injection allowed to infect 100% of the neurons in the VTA and SNc. This could for example explain in part the mismatch between the number of VTA dopamine neurons counted in panel 2G when comparing TH and mCherry counts. Also, I see that the mCherry counts were not provided at the 2-week time point. If the mCherry had been expressed genetically by crossing the DAT-Cre mice with a floxed fluorescent reported mice, the interpretation would have been simpler. In this context, I am not convinced of the benefit of the mCherry quantifications. The authors should consider either removing these results from the final manuscript or discussing this important limitation.

      Although the authors conclude that there is a global decrease in the number of dopamine neurons after 4 weeks of CNO treatment, the post-hoc tests failed to confirm that the decrease in dopamine number was significant in the SNc, the region most relevant to Parkinson's. This could be due to the fact that only a small number of mice were tested. A "n" of just 4 or 5 mice is very small for a stereological counting experiment. As such, this experiment was clearly underpowered at the statistical level. Also, the choice of the image used to illustrate this in panel 2G should be reconsidered: the image suggests that a very large loss of dopamine neurons occurred in the SNc and this is not what the numbers show. A more representative image should be used.

      In Figure 3, the authors attempt to compare intracellular calcium levels in dopamine neurons using GCaMP6 fluorescence. Because this calcium indicator is not quantitative (unlike ratiometric sensors such as Fura2), it is usually used to quantify relative changes in intracellular calcium. The present use of this probe to compare absolute values is unusual and the validity of this approach is unclear. This limitation needs to be discussed. The authors also need to refer in the text to the difference between panels D and E of this figure. It is surprising that the fluctuations in calcium levels were not quantified. I guess the hypothesis was that there should be more or larger fluctuations in the mice treated with CNO if the CNO treatment led to increased firing. This needs to be clarified.

      Although the spatial transcriptomic results are intriguing and certainly a great way to start thinking about how the CNO treatment could lead to the loss of dopamine neurons, the presented results, the focussing of some broad classes of differentially expressed genes and on some specific examples, do not really suggest any clear mechanism of neurodegeneration. It would perhaps be useful for the authors to use the obtained data to validate that a state of chronic depolarization was indeed induced by the chronic CNO treatment. Were genes classically linked to increased activity like cfos or bdnf elevated in the SNc or VTA dopamine neurons? In the striatum, the authors report that the levels of DARP32, a gene whose levels are linked to dopamine levels, are unchanged. Does this mean that there were no major changes in dopamine levels in the striatum of these mice?

      The usefulness of comparing the transcriptome of human PD SNc or VTA sections to that of the present mouse model should be better explained. In the human tissues, the transcriptome reflects the state of the tissue many years after extensive loss of dopamine neurons. It is expected that there will be few if any SNc neurons left in such sections. In comparison, the mice after 7 days of CNO treatment do not appear to have lost any dopamine neurons. As such, how can the two extremely different conditions be reasonably compared?

      Comments on the discussion:

      In the discussion, the authors state that their calcium photometry results support a central role of calcium in activity-induced neurodegeneration. This conclusion, although plausible because of the very broad pre-existing literature linking calcium elevation (such as in excitotoxicity) to neuronal loss, should be toned down a bit as no causal relationship was established in the experiments that were carried out in the present study.

      In the discussion, the authors discuss some of the parallel changes in gene expression detected in the mouse model and in the human tissues. Because few if any dopamine neurons are expected to remain in the SNc of the human tissues used, this sort of comparison has important conceptual limitations and these need to be clearly addressed.

      A major limitation of the present discussion is that it does not discuss the possibility that the observed phenotypes are caused by the induction of a chronic state of depolarization block by the chronic CNO treatment. I encourage the authors to consider and discuss this hypothesis. Also, the authors need to discuss the fact that previous work was only able to detect an increase in the firing rate of dopamine neurons after more than 95% loss of dopamine neurons. As such, the authors need to clearly discuss the relevance of the present model to PD. Are changes in firing rate a driver of neuronal loss in PD, as the authors try to make the case here, or are such changes only a secondary consequence of extensive neuronal loss (for example because a major loss of dopamine would lead to reduced D2 autoreceptor activation in the remaining neurons, and to reduced autoreceptor-mediated negative feedback on firing). This needs to be discussed.

      There is a very large, multi-decade literature on calcium elevation and its effects on neuronal loss in many different types of neurons. The authors should discuss their findings in this context and refer to some of this previous work. In a nutshell, the observations of the present manuscript could be summarized by stating that the chronic membrane depolarization induced by the CNO treatment is likely to induce a chronic elevation of intracellular calcium and this is then likely to activate some of the well-known calcium-dependent cell death mechanisms. Whether such cell death is linked in any way to PD is not really demonstrated by the present results.

      The authors are encouraged to perform a thorough revision of the discussion to address all of these issues, discuss the major limitations of the present model, and refer to the broad pre-existing literature linking membrane depolarization, calcium, and neuronal loss in many neuronal cell types.

    5. Author response:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, the authors investigated the effect of chronic activation of dopamine neurons using chemogenetics. Using Gq-DREADDs, the authors chronically activated midbrain dopamine neurons and observed that these neurons, particularly their axons, exhibit increased vulnerability and degeneration, resembling the pathological symptoms of Parkinson's disease. Baseline calcium levels in midbrain dopamine neurons were also significantly elevated following the chronic activation. Lastly, to identify cellular and circuit-level changes in response to dopaminergic neuronal degeneration caused by chronic activation, the authors employed spatial genomics (Visium) and revealed comprehensive changes in gene expression in the mouse model subjected to chronic activation. In conclusion, this study presents novel data on the consequences of chronic hyperactivation of midbrain dopamine neurons.

      Strengths:

      This study provides direct evidence that the chronic activation of dopamine neurons is toxic and gives rise to neurodegeneration. In addition, the authors achieved the chronic activation of dopamine neurons using water application of clozapine-N-oxide (CNO), a method not commonly employed by researchers. This approach may offer new insights into pathophysiological alterations of dopamine neurons in Parkinson's disease. The authors also utilized state-of-the-art spatial gene expression analysis, which can provide valuable information for other researchers studying dopamine neurons. Although the authors did not elucidate the mechanisms underlying dopaminergic neuronal and axonal death, they presented a substantial number of intriguing ideas in their discussion, which are worth further investigation.

      We thank the reviewer for these positive comments.

      Weaknesses:

      Many claims raised in this paper are only partially supported by the experimental results. So, additional data are necessary to strengthen the claims. The effects of chronic activation of dopamine neurons are intriguing; however, this paper does not go beyond reporting phenomena. It lacks a comprehensive explanation for the degeneration of dopamine neurons and their axons. While the authors proposed possible mechanisms for the degeneration in their discussion, such as differentially expressed genes, these remain experimentally unexplored.

      We thank the reviewer for this review. We do believe that the manuscript has a mechanistic component, as the central experiments involve direct manipulation of neuronal activity, and we show an increase in calcium levels and gene expression changes in dopamine neurons that coincide with the degeneration. However, we agree that deeper mechanistic investigation would strengthen the conclusions of the paper. We have planned several important revisions, including the addition of CNO behavioral controls, manipulation of intracellular calcium using isradipine, additional transcriptomics experiments and further validation of findings. We anticipate that these additions will significantly bolster the conclusions of the paper.

      Reviewer #2 (Public Review):

      Summary:

      Rademacher et al. present a paper showing that chronic chemogenetic excitation of dopaminergic neurons in the mouse midbrain results in differential degeneration of axons and somas across distinct regions (SNc vs VTA). These findings are important. This mouse model also has the advantage of showing a axon-first degeneration over an experimentally-useful time course (2-4 weeks). 2. The findings that direct excitation of dopaminergic neurons causes differential degeneration sheds light on the mechanisms of dopaminergic neuron selective vulnerability. The evidence that activation of dopaminergic neurons causes degeneration and alters mRNA expression is convincing, as the authors use both vehicle and CNO control groups, but the evidence that chronic dopaminergic activation alters circadian rhythm and motor behavior is incomplete as the authors did not run a CNO-control condition in these experiments.

      Strengths:

      This is an exciting and important paper.

      The paper compares mouse transcriptomics with human patient data.

      It shows that selective degeneration can occur across the midbrain dopaminergic neurons even in the absence of a genetic, prion, or toxin neurodegeneration mechanism.

      We thank the reviewer for these insightful comments.

      Weaknesses:

      Major concerns:

      (1) The lack of a CNO-positive, DREADD-negative control group in the behavioral experiments is the main limitation in interpreting the behavioral data. Without knowing whether CNO on its own has an impact on circadian rhythm or motor activity, the certainty that dopaminergic hyperactivity is causing these effects is lacking.

      This is an important point. Although we show that CNO does not produce degeneration of DA neuron terminals, we do not exclude a contribution to the behavioral changes. We agree that this behavioral control is necessary, and will address it in revision with a CNO-only running wheel cohort.

      (2) One of the most exciting things about this paper is that the SNc degenerates more strongly than the VTA when both regions are, in theory, excited to the same extent. However, it is not perfectly clear that both regions respond to CNO to the same extent. The electrophysiological data showing CNO responsiveness is only conducted in the SNc. If the VTA response is significantly reduced vs the SNc response, then the selectivity of the SNc degeneration could just be because the SNc was more hyperactive than the VTA. Electrophysiology experiments comparing the VTA and SNc response to CNO could support the idea that the SNc has substantial intrinsic vulnerability factors compared to the VTA.

      We agree that additional electrophysiology conducted in the VTA dopamine neurons would meaningfully add to our understanding of the selective vulnerability in this model, and will complete these experiments in revision.

      (3) The mice have access to a running wheel for the circadian rhythm experiments. Running has been shown to alter the dopaminergic system (Bastioli et al., 2022) and so the authors should clarify whether the histology, electrophysiology, fiber photometry, and transcriptomics data are conducted on mice that have been running or sedentary.

      We will explicitly clarify which mice had access to a running wheel in our revision. Briefly, mice for histology, electrophysiology, and transcriptomics all had access to a running wheel during their treatment. The mice used for photometry underwent about 7 days of running wheel access approximately 3 weeks prior to the beginning of the experiment. The photometry headcaps sterically prevented mice from having access to a running wheel in their home cage.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Rademacher and colleagues examined the effect on the integrity of the dopamine system in mice of chronically stimulating dopamine neurons using a chemogenetic approach. They find that one to two weeks of constant exposure to the chemogenetic activator CNO leads to a decrease in the density of tyrosine hydroxylase staining in striatal brain sections and to a small reduction of the global population of tyrosine hydroxylase positive neurons in the ventral midbrain. They also report alterations in gene expression in both regions using a spatial transcriptomics approach. Globally, the work is well done and valuable and some of the conclusions are interesting. However, the conceptual advance is perhaps a bit limited in the sense that there is extensive previous work in the literature showing that excessive depolarization of multiple types of neurons associated with intracellular calcium elevations promotes neuronal degeneration. The present work adds to this by showing evidence of a similar phenomenon in dopamine neurons.

      We thank the reviewer for the careful and thoughtful review of our manuscript.

      While extensive depolarization and associated intracellular calcium elevations promotes degeneration generally, we emphasize that the process we describe is novel. Indeed, prior studies delivering chronic DREADDs to vulnerable neurons in models of Alzheimer’s disease did not report an increase in neurodegeneration, despite seeing changes in protein aggregation (e.g. Yuan and Grutzendler, J Neurosci 2016, PMID: 26758850; Hussaini et al., PLOS Bio 2020, PMID: 32822389). Further, a critical finding from our study is that in our paradigm, this stressor does not impact all dopamine neurons equally, as the SNc DA neurons are more vulnerable than the VTA, mirroring selective vulnerability characteristic of Parkinson’s disease. This is consistent with a large body of literature that SNc dopamine neurons are less capable of handling large energetic and calcium loads compared to neighboring VTA neurons, and the finding that chronically altered activity is sufficient to drive this preferential loss is novel.

      In addition, we are not aware of prior studies that have chronically activated DREADDs to produce neurodegeneration. Other studies have shown that acute excitotoxic stressors can produce neuronal degeneration, but the chronic increase in activity is central to our approach.

      In terms of the mechanisms explaining the neuronal loss observed after 2 to 4 weeks of chemogenetic activation, it would be important to consider that dopamine neurons are known from a lot of previous literature to undergo a decrease in firing through a depolarization-block mechanism when chronically depolarized. Is it possible that such a phenomenon explains much of the results observed in the present study? It would be important to consider this in the manuscript.

      As discussed in greater detail in the results section below, our data suggests this may not be a prominent feature in our model. However, we cannot rule out a contribution of depolarization block, and will expand on the discussion of this possibility in the revised manuscript.

      The relevance to Parkinson's disease (PD) is also not totally clear because there is not a lot of previous solid evidence showing that the firing of dopamine neurons is increased in PD, either in human subjects or in mouse models of the disease. As such, it is not clear if the present work is really modelling something that could happen in PD in humans.

      We completely agree that evidence of increased dopamine neuron activity from human PD patients is lacking and the existing data are difficult to interpret without human controls. However, as we outline in the manuscript, multiple lines of evidence suggest that the activity level of dopamine neurons almost certainly does change in PD. Therefore, it is very important that we understand how changes in the level of neural activity influence the degeneration of DA neurons. In this paper we examine the impact of increased activity. Increased activity may be compensatory after initial dopamine neuron loss, or may be an initial driver of death (Rademacher & Nakamura, Exp Neurol 2024, PMID: 38092187). Beyond what is already discussed in the manuscript, additional support for increased activity in PD models include:

      - Elevated firing rates in asymptomatic MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488)

      - Increased frequency of spontaneous firing in patient-derived iPSC dopamine neurons and primary mouse dopamine neurons that overexpress synuclein (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060)

      - Increased spontaneous firing in dopamine neurons of rats injected with synuclein preformed fibrils compared to sham (Tozzi et al., Brain 2021, PMID: 34297092)

      We will include and further discuss these important examples in our revision.

      Similarly, in future studies, it will also be important to study the impact of decreasing DA neuron activity. There will be additional levels of complexity to accurately model changes in PD, which may differ between subtypes of the disease, the disease stage, and the subtype of dopamine neuron. Our study models the possibility of chronically increased pacemaking, and interpretation of our results will be informed as we learn more about how the activity of DA neurons changes in humans in PD. We will discuss and elaborate on these important points in the revision.

      Comments on the introduction:

      The introduction cites a 1990 paper from the lab of Anthony Grace as support of the fact that DA neurons increase their firing rate in PD models. However, in this 1990 paper, the authors stated that: "With respect to DA cell activity, depletions of up to 96% of striatal DA did not result in substantial alterations in the proportion of DA neurons active, their mean firing rate, or their firing pattern. Increases in these parameters only occurred when striatal DA depletions exceeded 96%." Such results argue that an increase in firing rate is most likely to be a consequence of the almost complete loss of dopamine neurons rather than an initial driver of neuronal loss. The present introduction would thus benefit from being revised to clarify the overriding hypothesis and rationale in relation to PD and better represent the findings of the paper by Hollerman and Grace.

      We agree that the findings of Hollerman and Grace support compensatory changes in dopamine neuron activity in response to loss of dopamine neurons, rather than informing whether dopamine neuron loss can also be an initial driver of activity. We will clarify this point in our revision. In addition, the results of other studies on this point are mixed: a 50% reduction in dopamine neurons didn’t alter firing rate or bursting (Harden and Grace, J Neurosci 1995, PMID: 7666198; Bilbao et al, Brain Res 2006, PMID: 16574080), while a 40% loss was found to increase firing rate and bursting (Chen et al, Brain Res 2009. PMID: 19545547) and larger reductions alter burst firing (Hollerman & Grace, Brain Res 1990, PMID: 2126975; Stachowiak et al, J Neurosci 1987, PMID: 3110381). Importantly, even if compensatory, such late-stage increases in dopamine neuron activity may contribute to disease progression and drive a vicious cycle of degeneration in surviving neurons. In addition, we also don’t know how the threshold of dopamine neuron loss and altered activity may differ between mice and humans, and PD patients do not present with clinical symptoms until ~30-60% of nigral neurons are lost (Burke & O’Malley, Exp Neurol 2013, PMID: 22285449; Shulman et al, Annu Rev Pathol 2011, PMID: 21034221).

      Other lines of evidence support the potential role of hyperactivity in disease initiation, including increased activity before dopamine neuron loss in MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488), increased spontaneous firing in patient-derived iPSC dopamine neurons (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060), and increased activity observed in genetic models of PD (Bishop et al., J Neurophysiol 2010, PMID: 20926611; Regoni et al., Cell Death Dis 2020,  PMID: 33173027).

      It would be good that the introduction refers to some of the literature on the links between excessive neuronal activity, calcium, and neurodegeneration. There is a large literature on this and referring to it would help frame the work and its novelty in a broader context.

      We agree that a discussion of hyperactivity, calcium, and neurodegeneration would benefit the introduction. While we briefly discuss calcium and neurodegeneration in the discussion, we will expand on this literature in both the introduction and discussion sections. We will carefully review and contextualize our work within existing frameworks of calcium and neurodegeneration (e.g. Surmeier & Schumacker, J Biol Chem 2013, PMID: 23086948; Verma et al., Transl Neurodegener 2022, PMID: 35078537). We believe that the novelty of our study lies in 1) a chronic chemogenetic activation paradigm via drinking water, 2) demonstrating selective vulnerability of dopamine neurons as a result of altering their activity/excitability alone, and 3) comparing mouse and human spatial transcriptomics.

      Comments on the results section:

      The running wheel results of Figure 1 suggest that the CNO treatment caused a brief increase in running on the first day after which there was a strong decrease during the subsequent days in the active phase. This observation is also in line with the appearance of a depolarization block.

      The authors examined many basic electrophysiological parameters of recorded dopamine neurons in acute brain slices. However, it is surprising that they did not report the resting membrane potential, or the input resistance. It would be important that this be added because these two parameters provide key information on the basal excitability of the recorded neurons. They would also allow us to obtain insight into the possibility that the neurons are chronically depolarized and thus in depolarization block.

      We do report the input resistance in Supplemental Figure 1C, which was unchanged in CNO-treated animals compared to controls. We did not report the resting membrane potential because many of the DA neurons were spontaneously firing. However, we will report the initial membrane potential on first breaking into the cell for the whole cell recordings in the revision, which did not vary between groups. This is still influenced by action potential activity, but is the timepoint in the recording least impacted by dialyzing of the neuron by the internal solution. We observed increased spontaneous action potential activity ex vivo in slices from CNO-treated mice (Figure 1D), thus at least under these conditions these dopamine neurons are not in depolarization block. We also did not see strong evidence of changes in other intrinsic properties of the neurons with whole cell recordings (e.g. Figure S1C). Overall, our electrophysiology experiments are not consistent with the depolarization block model, at least not due to changes in the intrinsic properties of the neurons. Although our ex vivo findings cannot exclude a contribution of depolarization block in vivo, we do show that CNO-treated mice removed from their cages for open field testing continue to have a strong trend for increased activity for approximately 10 days (S1E).  This finding is also consistent with increased activity of the DA neurons. We will add discussion of these important considerations in the revision.

      It is great that the authors quantified not only TH levels but also the levels of mCherry, co-expressed with the chemogenetic receptor. This could in principle help to distinguish between TH downregulation and true loss of dopamine neuron cell bodies. However, the approach used here has a major caveat in that the number of mCherry-positive dopamine neurons depends on the proportion of dopamine neurons that were infected and expressed the DREADD and this could very well vary between different mice. It is very unlikely that the virus injection allowed to infect 100% of the neurons in the VTA and SNc. This could for example explain in part the mismatch between the number of VTA dopamine neurons counted in panel 2G when comparing TH and mCherry counts. Also, I see that the mCherry counts were not provided at the 2-week time point. If the mCherry had been expressed genetically by crossing the DAT-Cre mice with a floxed fluorescent reported mice, the interpretation would have been simpler. In this context, I am not convinced of the benefit of the mCherry quantifications. The authors should consider either removing these results from the final manuscript or discussing this important limitation.

      We thank the reviewer for this insightful comment, and we agree that this is a caveat of our mCherry quantification. Quantitation of the number of mCherry+ DA neurons specifically informs the impact on transduced DA neurons, and mCherry appears to be less susceptible to downregulation versus TH. As the reviewer points out, it carries the caveat that there is some variability between injections. Nonetheless, we believe that it conveys useful complementary data. As suggested, we will discuss this caveat in our revision. Note that mCherry was not quantified at the two-week timepoint because there is no loss of TH+ cells at that time.

      Although the authors conclude that there is a global decrease in the number of dopamine neurons after 4 weeks of CNO treatment, the post-hoc tests failed to confirm that the decrease in dopamine number was significant in the SNc, the region most relevant to Parkinson's. This could be due to the fact that only a small number of mice were tested. A "n" of just 4 or 5 mice is very small for a stereological counting experiment. As such, this experiment was clearly underpowered at the statistical level. Also, the choice of the image used to illustrate this in panel 2G should be reconsidered: the image suggests that a very large loss of dopamine neurons occurred in the SNc and this is not what the numbers show. A more representative image should be used.

      We agree that the stereology experiments were performed on relatively small numbers of animals. Combined with the small effect size, this may have contributed to the post-hoc tests showing a trend of p=0.1 for both the TH and mCherry dopamine cell counts in the SN at 4 weeks. As part of the planned experiments for our revision, we will perform an additional stereologic analysis to further assess the loss of SNc dopamine neurons. We will also review and ensure the images are representative.

      In Figure 3, the authors attempt to compare intracellular calcium levels in dopamine neurons using GCaMP6 fluorescence. Because this calcium indicator is not quantitative (unlike ratiometric sensors such as Fura2), it is usually used to quantify relative changes in intracellular calcium. The present use of this probe to compare absolute values is unusual and the validity of this approach is unclear. This limitation needs to be discussed. The authors also need to refer in the text to the difference between panels D and E of this figure. It is surprising that the fluctuations in calcium levels were not quantified. I guess the hypothesis was that there should be more or larger fluctuations in the mice treated with CNO if the CNO treatment led to increased firing. This needs to be clarified.

      We thank the reviewer for this comment. We understand that this method of comparing absolute values is unconventional. However, these animals were tested concurrently on the same system, and a clear effect on the absolute baseline was observed. We will include a caveat of this in our discussion. Panel D of this figure shows the raw, uncorrected photometry traces, whereas panel E shows the isosbestic corrected traces for the same recording. In panel E, the traces follow time in ascending order. We will also include frequency and amplitude data for these recordings.   

      Although the spatial transcriptomic results are intriguing and certainly a great way to start thinking about how the CNO treatment could lead to the loss of dopamine neurons, the presented results, the focusing of some broad classes of differentially expressed genes and on some specific examples, do not really suggest any clear mechanism of neurodegeneration. It would perhaps be useful for the authors to use the obtained data to validate that a state of chronic depolarization was indeed induced by the chronic CNO treatment. Were genes classically linked to increased activity like cfos or bdnf elevated in the SNc or VTA dopamine neurons? In the striatum, the authors report that the levels of DARP32, a gene whose levels are linked to dopamine levels, are unchanged. Does this mean that there were no major changes in dopamine levels in the striatum of these mice?

      We will review the expression of activity-related genes in our dataset, although we must keep in mind that these genes may behave differently in the context of chronic activation as opposed to acutely increased activity. We will also include experiments assessing striatal dopamine levels by HPLC in the revision.

      The usefulness of comparing the transcriptome of human PD SNc or VTA sections to that of the present mouse model should be better explained. In the human tissues, the transcriptome reflects the state of the tissue many years after extensive loss of dopamine neurons. It is expected that there will be few if any SNc neurons left in such sections. In comparison, the mice after 7 days of CNO treatment do not appear to have lost any dopamine neurons. As such, how can the two extremely different conditions be reasonably compared?

      Our mouse model and human PD progress over distinct timescales, as is the case with essentially all mouse models of neurodegenerative diseases. Nonetheless, in our view there is still great value in comparing gene expression changes in mouse models with those in human disease. It seems very likely that the same pathologic processes that drive degeneration early in the disease continue to drive degeneration later in the disease. Note that we have tried to address the discrepancy in time scales in part by comparing to early PD samples when there is more limited SNc DA neuron loss. Please note the numbers of DA neurons within the areas we have selected for sampling (Figure at right). Therefore, we can indeed use spatial transcriptomics to compare dopamine neurons from mice with initial degeneration and patients where degeneration is ongoing during their disease.

      Author response image 1.

      Violin plot of DA neuron proportions sampled within the vulnerable SNV (deconvoluted RCTD method used in unmasked tissue sections of the SNV).

      Control and early PD subjects.

      Comments on the discussion:

      In the discussion, the authors state that their calcium photometry results support a central role of calcium in activity-induced neurodegeneration. This conclusion, although plausible because of the very broad pre-existing literature linking calcium elevation (such as in excitotoxicity) to neuronal loss, should be toned down a bit as no causal relationship was established in the experiments that were carried out in the present study.

      Our model utilizes hM3Dq-DREADDs that function by increasing intracellular calcium to increase neuronal excitability, and our results show increased Ca2+ by fiber photometry and changes to Ca2+-related genes, strongly suggesting a causal relation and crucial role of calcium in the mechanism of degeneration. However, we agree that we have not experimentally proven this point, as we acknowledged in the text. Additionally, we have planned revision experiments involving chronic isradipine treatment to further test the role of calcium in the mechanism of degeneration in this model.

      In the discussion, the authors discuss some of the parallel changes in gene expression detected in the mouse model and in the human tissues. Because few if any dopamine neurons are expected to remain in the SNc of the human tissues used, this sort of comparison has important conceptual limitations and these need to be clearly addressed.

      As discussed, we can sample SN DA neurons in early PD (see figure above), and in our view there is great value for such comparisons. We agree that discussion of appropriate caveats is warranted and this will be clearly addressed in the revision.

      A major limitation of the present discussion is that it does not discuss the possibility that the observed phenotypes are caused by the induction of a chronic state of depolarization block by the chronic CNO treatment. I encourage the authors to consider and discuss this hypothesis.

      As discussed above, our analyses of DA neuron firing in slices and open field testing to date do not support a prominent contribution of depolarization block with chronic CNO treatment. However, we cannot rule out this hypothesis, therefore we will include additional electrophysiology experiments and add discussion of this important consideration.  

      Also, the authors need to discuss the fact that previous work was only able to detect an increase in the firing rate of dopamine neurons after more than 95% loss of dopamine neurons. As such, the authors need to clearly discuss the relevance of the present model to PD. Are changes in firing rate a driver of neuronal loss in PD, as the authors try to make the case here, or are such changes only a secondary consequence of extensive neuronal loss (for example because a major loss of dopamine would lead to reduced D2 autoreceptor activation in the remaining neurons, and to reduced autoreceptor-mediated negative feedback on firing). This needs to be discussed.

      As discussed above, while increases in dopamine neuron activity may be compensatory after loss of neurons, the precise percentage required to induce such compensatory changes is not defined in mice and varies between paradigms, and the threshold level is not known in humans. We also reiterate that a compensatory increase in activity could still promote the degeneration of critical surviving DA neurons, whose loss underlies the substantial decline in motor function that typically occurs over the course of PD. Moreover, there are also multiple lines of evidence to suggest that changes in activity can initiate and drive dopamine neuron degeneration (Rademacher & Nakamura, Exp Neurol 2024). For example, overexpression of synuclein can increase firing in cultured dopamine neurons (Dagra et al., NPJ Parkinsons Dis 2021, PMID: 34408150) while mice expressing mutant Parkin have higher mean firing rates (Regoni et al., Cell Death Dis 2020,  PMID: 33173027). Similarly, an increased firing rate has been reported in the MitoPark mouse model of PD at a time preceding DA neuron degeneration (Good et al., FASEB J 2011, PMID: 21233488). We also acknowledge that alterations to dopamine neuron activity are likely complex in PD, and that dopamine neuron health and function can be impacted not just by simple increases in activity, but also by changes in activity patterns and regularity. We will amend our discussion to include the important caveat of changes in activity occurring as compensation, as well as further evidence of changes in activity preceding dopamine neuron death.

      There is a very large, multi-decade literature on calcium elevation and its effects on neuronal loss in many different types of neurons. The authors should discuss their findings in this context and refer to some of this previous work. In a nutshell, the observations of the present manuscript could be summarized by stating that the chronic membrane depolarization induced by the CNO treatment is likely to induce a chronic elevation of intracellular calcium and this is then likely to activate some of the well-known calcium-dependent cell death mechanisms. Whether such cell death is linked in any way to PD is not really demonstrated by the present results. The authors are encouraged to perform a thorough revision of the discussion to address all of these issues, discuss the major limitations of the present model, and refer to the broad pre-existing literature linking membrane depolarization, calcium, and neuronal loss in many neuronal cell types.

      While our model demonstrates classic excitotoxic cell death pathways, we would like to emphasize both the chronic nature of our manipulation and the progressive changes observed, with increasing degeneration seen at 1, 2, and 4 weeks of hyperactivity in an axon-first manner. This is a unique aspect of our study, in contrast to much of the previous literature which has focused on shorter timescales. Thus, while we will revise the discussion to more comprehensively acknowledge previous studies of calcium-dependent neuron cell death, we believe we have made several new contributions that are not predicted by existing literature. We have shown that this chronic manipulation is specifically toxic to nigral dopamine neurons, and the data that VTA dopamine neurons continue to be resilient even at 4 weeks is interesting and disease-relevant. We therefore do not want to use findings from other neuron types to draw assumptions about DA neurons, which are a unique and very diverse population. We acknowledge that as with all preclinical models of PD, we cannot draw definitive conclusions about PD with this data. However, we reiterate that we strongly believe that drawing connections to human disease is important, as dopamine neuron activity is very likely altered in PD and a clearer understanding of how dopamine neuron survival is impacted by activity will provide insight into the mechanisms of PD.

    1. eLife Assessment

      This study demonstrates that endothelial toll-like receptor 4 is a central regulator of leptomeningeal inflammation in neonatal E. coli meningitis. The data are derived from cell-type-specific gene knockouts in mice and cultured endothelial cells and are convincing. This work is important as it advances our understanding of host cellular processes and molecular pathways underlying meningitis pathogenesis and specifically expands the knowledge of how breakdown of the blood brain barrier contributes to the pathogenesis.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Seegren and colleagues demonstrate that in a mouse model of neonatal E. coli meningitis, loss of toll-like receptor 4 (TLR4) in VE-cadherin+ endothelial cells and a subset of meningeal fibroblasts leads to a marked decrease in transcriptional dysregulation across multiple leptomeningeal cell types, a decrease in vascular permeability, and a decrease in macrophage abundance. In contrast, loss of macrophage TLR4 had less pronounced effects. Using cultured wildtype and TLR4-knockout endothelial cells, the authors further demonstrate that TLR4 signaling leads to reversible internalization of the tight junction protein claudin-5, establishing a potential mechanism of increased vascular permeability. Authors also show that claudin-5 internalization is independent of NF-κB. Finally, the authors use RNA-sequencing of wildtype and TLR4-knockout endothelial cells to define the TLR4-dependent cell-autonomous transcriptional response to E. coli.

      Comments on revised version.

      The authors have considerably improved and strengthened the work through the addition of new experimental data, new data analyses, and modifications to their interpretation. Notably, the authors used additional Cre-reporter mice to clarify that Cdh5-CreER is active in endothelial cells and some meningeal fibroblasts, and thus revised nomenclature and interpretation to acknowledge that the Tlr4fl/-;Cdh5-CreER cKO (Tlr4-VEKO) is not exclusively endothelial. The authors also demonstrated that Tlr4-VEKO does not affect peripheral E.coli burden, but acknowledge that changes to periphery-derived signals (e.g., cytokines) may contribute to observed leptomeningeal phenotypes.

      The authors added PCA plots to show similarity in gene expression shifts across biological replicates (mice). This provides support for the claim that Tlr4-VEKO attenuates infection-associated transcriptional changes. With respect to differential expression analysis, I agree with authors that characteristics of individual cells (e.g. heterogeneity) are of interest. I remain concerned, however, that the formal differential analysis strategy appears to consider cells as independent experimental units, which they are not because a single cell cannot be randomly assigned to an experimental group (control or cKO, uninfected or infected). The mouse is the correct experimental unit for a comparison across these groups because it can be randomized. I appreciate that many of the gene expression changes appear consistent across mice (e.g. Figure 1 - Figure supplement 7) and that there are clear infection- and genotype-associated phenotypes in other assays. I would simply caution that the authors' analysis strategy likely leads to a larger number of type I errors (false positives) than is generally accepted; a mixed (hierarchical) model or pseudo-bulk approach would be more appropriate for future studies.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use a postnatal mouse model of E. coli bacterial meningitis and a mouse brain endothelioma cell line combined with cell type specific gene deletion to study the function of endothelial TLR4, a cell surface receptor that recognizes gram positive bacterial wall components, in the local leptomeningeal (LPM) response with a focus on endothelial barrier breakdown mediated by TLR4. Single cell transcriptional profiling and imaging studies using wholemount preps of the LPM support that LPM endothelial, CD206+ local macrophage and LPM fibroblast and arachnoid barrier cell inflammatory response and is abrogated in endothelial specific KO of TLR4, pointing to a role for endothelial TLR4 in local LPM response. Culture studies using Bend3.1 cells (a mouse brain endothelioma cell line) support a direct role for TLR4 in the bacteria-mediated inflammatory response and in internalization of Cldn5 via the endosomal-lysosomal pathway, resulting in loss of barrier integrity

      Strengths:

      The local LPM cell response in meningitis and the role of specific LPM cells in inflammation and CNS barrier breakdown has not been extensively studied, despite ample evidence for primary immune response in the meninges in human patients and in animal models. The authors employ a robust, multi-model approach using both in vivo and in vitro models with cell-type specific knockout to study the function of TLR4 in brain endothelial cell response. The authors nicely combine functional barrier assays with IF for junctional localization in their experimental design and they delve into potential mechanisms of Cldn5 internalization using markers of endosomal-lysomal pathway localization. The authors also describe a new type of barrier assay using a streptavidin-coated plates upon which barrier forming cell cultures can be plated, this could be a very useful alternative or complement to other size-selective barrier assays and presumably could work for other barrier forming cell types, like epithelial cells.

      Comments on revised version.

      In their revision, the authors addressed prior noted weaknesses with new data and analysis. They now show that TLR4-VE-cad cKO mice have a largely similar disease progression as control mice, including increased bacterial burden in the LPM and brain. This underscores that that the reduced vascular leakage and blunted inflammatory response is due to loss of TLR4 response to bacteria on VE-cad recombined cells and not because the mice are protected from meningitis. The authors also performed additional experiments to show that Cldn5 internalization via the endosomal-lysosomal pathway is independent of NFKB signaling. The authors also added in important discussion points about how their results fit into the broader literature on TLR4 in BBB endothelial cell junctional protein localization and prior work on meningitis in global TLR4.

    4. Reviewer #3 (Public review):

      Summary:

      This study investigates the molecular underpinnings of immune responses in the leptomeninges in neonatal bacterial meningitis. Bacterial meningitis is a major disease burden, particularly for neonates, and it has previously been noted that the meningeal immune environment in infants is permissive to opportunistic infection (Kim et al., Sci Immunol, 2023). There is less known about the contribution of the stromal compartment to meningeal immune responses. Seegren et al. interrogate the role of leptomeningeal endothelium in host defense in E. coli infected neonatal mice using mouse genetic tools to delete the LPS receptor Tlr4 from either endothelial cells/stromal cells (using Cdh5-CreER) or myeloid cells (using LysM-Cre). The authors use snRNAseq, cleared cortical mounts, and in vitro work to define the impact of E. coli infection on leptomeningeal endothelial cells. This study uses a range of innovative techniques to probe the role of the stromal compartment in meningitis. With additional experiments to confirm the specificity of their Cre models, this strengthens the interpretation of the study significantly. The only major weakness is the inability to confirm TLR4 knockout in myeloid cells.

      Strengths:

      This study makes excellent use of cleared cortical mounts to examine the biology of the leptomeninges, in particular, changes to the endothelium, with unprecedented detail. In combination with high-quality sequencing data provide new insights into the impact of meningitis on the leptomeninges. The data presented by the authors is of very high quality.

      The authors have also done substantial work to address my two major comments regarding 1) the specificity of their Cre systems and 2) peripheral impacts of the interventions.

      (1) The authors identified and acknowledged some impacts in the leptomeningeal stroma (the relatively high level of recombination in ECs vs FBs presumably reflects a single low dose being given, where other groups have done more aggressive tamoxifen regimens that drive recombination in FBs as well). Given the incomplete recombination in the leptomeningeal FBs, I agree with their conclusion that it is probably endothelial driven. Acknowledging the contributions of other myeloid cells with the L. The Cre-NLS experiments with nuclear markers provided excellent data and had beautiful staining.

      (2) The authors did not observe differences in bacterial burden in peripheral organs in either CKO model, suggesting that CNS impacts are not downstream of peripheral bacterial control.

      Weaknesses:

      (1) The inducible Cre lines used by the authors target peripheral tissues as well as CNS tissues. Although this is mollified by the lack of impact on peripheral disease burden.

      (2) The authors were not able to confirm TLR4 knockout in myeloid cells, and this caveat is acknowledged. The lack of response in TLR4 VEKO mice strongly suggests successful conditional knockout.

      (3) The cell line model (bEnd.3) is a relatively low fidelity model of BBB endothelial cells. The authors acknowledge this, and it is likely that endothelial cell responses to LPS are highly conserved.

      (4) It is perhaps not surprising that Tlr4 is required for meningitis responses with E. coli. However, it is unclear if these findings can be generalised to other, more common, meningitis infections (streptococcal/pneumococcal).

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Seegren and colleagues demonstrate that in a mouse model of neonatal E. coli meningitis, loss of endothelial toll-like receptor 4 (TLR4) leads to a marked decrease in transcriptional dysregulation across multiple leptomeningeal cell types, a decrease in vascular permeability, and a decrease in macrophage abundance. In contrast, loss of macrophage TLR4 had less pronounced effects. Using cultured wild-type and TLR4knockout endothelial cells, the authors further demonstrate that TLR4-NF-κB signaling leads to reversible internalization of the tight junction protein claudin-5, establishing a potential mechanism of increased vascular permeability. Finally, the authors use RNA sequencing of wild-type and TLR4-knockout endothelial cells to define the TLR4dependent cell-autonomous transcriptional response to E. coli.

      Strengths:

      (1) The authors address an important, well-motivated hypothesis related to the cellular and molecular mechanisms of leptomeningeal inflammation.

      (2) The authors use model systems (mouse conditional knockouts and cultured endothelial cells) that are appropriate to address their hypotheses. The data are of high quality.

      Weaknesses:

      (1) The authors perform single-nucleus RNA-seq on dissected leptomeninges from control and E. coli-infected mice across three genotypes (WT, Tlr4MKO, and Tlr4ECKO). A major discovery from this experiment, as summarized by the authors, is: "Tlr4ECKO mice exhibited a global attenuation of infection-induced transcriptional responses across all major leptomeningeal cell types, as judged by the positions of cell clusters in the UMAP." This conclusion could be considerably strengthened by improving the qualitative and quantitative analysis.

      Thank you for this comment. We agree that the UMAP-based interpretation would benefit from additional qualitative and quantitative support. We have expanded the snRNA-seq analysis with additional images and supplemental figures (Figure 1 – figure supplement 3, Figure 1 – figure supplement 5, and Figure 1 – figure supplement 6). The first and third of these new supplemental figures show dot plots for each major leptomeningeal cell type, for each genotype, for the two experimental conditions (infected vs. uninfected), and for individual genes in three immune-related gene sets (NF-kB and TNF-α, JAK-STAT, and IFN-ɣ), providing a more explicit comparison of infection-induced transcriptional responses across genotypes. The second of these new supplemental figure shows principal component analysis (PCA) of the individual snRNA-seq datasets (one mouse per dataset) for each genotype and experimental condition, demonstrating that the observed transcriptional shifts are consistent across biological replicates. Finally, Figure 1 – figure supplement 7, which was included in the original submission, shows changes in the most up- and down-regulated genes (based on adjusted p-value or fold change) in endothelial and myeloid cells across individual mice and genotypes/conditions, further supporting the genotype-dependent effects at the level of individual animals.

      (2) The authors interpret E. coli infection-induced increases in leptomeningeal sulfo-NHSbiotin as evidence of compromised BBB integrity (i.e., extravasation from the vasculature) (Results, page 7), but another possible route in this context is sulfo-NHS-biotin entry from the dura across a compromised arachnoid barrier. The complete rescue in Tlr4ECKOs is strongly suggestive that the vascular route dominates, but it would strengthen the work if the authors could assess arachnoid barrier fidelity (e.g. via immunohistochemistry). At a minimum, authors should mention that the sulfo-NHS-biotin signal in this context may represent both vascular and arachnoid barrier extravasation.

      Thank you for this comment. We agree that our data cannot rule out leakage across the arachnoid barrier during infection. While the rescue observed in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice strongly supports a dominant vascular contribution, we acknowledge that the sulfo-NHS-biotin signal may reflect permeability at both the vascular and arachnoid barriers. We do not think there is a clear way to directly test this possibility functionally, since the arachnoid barrier appears intact by confocal microscopy. Subtle differences in barrier cell morphology might be detectable by electron microscopy, but this would not definitively address whether infection permits molecular passage across the arachnoid barrier. We have followed the reviewer’s suggestion and revised the Results section to reflect this interpretation. Specifically, we added the following: “The simplest interpretation of these data is that the site of sulfo-NHS biotin leakage is primarily vascular. However, we cannot exclude some contribution from increased arachnoid barrier permeability.”

      (3) The authors state that "deletion of TLR4 prevented both NF-κB nuclear translocation and Cldn5 internalization in response to E. coli (Figure 4A-D)" (Results, page 9). In Figures 4C and D, however, there is no indicator of a statistical test directly comparing the two genotypes. A comparison of within-genotype P-values should not be used to support a genotype difference (PMID: 34726155).

      Thank you for pointing out this omission. We have updated the figures so that the between-genotype p-values are shown for those panels (including this panel) that had not previously shown them.

      (4) In the first paragraph of the Results, the authors summarize the meningeal layers as (1) pia, (2) subarachnoid space, (3) arachnoid, and (4) dura, and then state "The second and third layers constitute the leptomeninges." This definition of leptomeninges seems to omit the pia, which is widely considered part of the leptomeninges (PMID: 37776854).

      Thank you for pointing out this error, which has now been corrected.

      (5) The Cdh5-CreER/+;Tlr4 fl/- mouse lacks TLR4 in all endothelial cells (i.e., in peripheral organs as well as CNS/leptomeninges), and, as the authors note, the periphery is exposed to E. coli. It would be helpful if the authors could comment in the Discussion on the possibility that peripheral effects (e.g., peripheral endothelial cytokine production, changes to blood composition as a result of changes to peripheral endothelial permeability) may contribute to the observed leptomeningeal phenotypes.

      Thank you for raising this point. We agree that peripheral responses could contribute to the observed leptomeningeal phenotypes in this model. We have added two sentences to the second paragraph of the Discussion to address this: “We note that these experiments do not distinguish between local vs. distal anatomic sources of LPS or downstream effector molecules, such as cytokines, that activate the leptomeningeal inflammatory response (Huang et al., 2021). Histologic observations of RFP-expressing E. coli in the brain, liver, and lungs, together with positive blood cultures, indicate substantial systemic dissemination in this model. Thus, the inflammatory responses of leptomeningeal cells likely reflect exposure to bacterial products and inflammatory mediators derived from both local meningeal and peripheral sources.”

      Reviewer #2 (Public review):

      Summary:

      The authors use a postnatal mouse model of E. coli bacterial meningitis and a mouse brain endothelioma cell line combined with cell-type-specific gene deletion to study the function of endothelial TLR4, a cell surface receptor that recognizes gram positive bacterial wall components, in the local leptomeningeal (LPM) response with a focus on endothelial barrier breakdown mediated by TLR4. Single-cell transcriptional profiling and imaging studies using whole-mount preps of the LPM support that LPM endothelial, CD206+ local macrophage and LPM fibroblast and arachnoid barrier cell inflammatory response and is abrogated in endothelial-specific KO of TLR4, pointing to a role for endothelial TLR4 in local LPM response. Culture studies using Bend3.1 cells (a mouse brain endothelioma cell line) support a direct role for TLR4 in the bacteria-mediated inflammatory response and in internalization of Cldn5 via the endosomal-lysosomal pathway, resulting in loss of barrier integrity

      Strengths:

      The local LPM cell response in meningitis and the role of specific LPM cells in inflammation and CNS barrier breakdown have not been extensively studied, despite ample evidence for primary immune response in the meninges in human patients and in animal models. The authors employ a robust, multi-model approach using both in vivo and in vitro models with cell-type-specific knockout to study the function of TLR4 in brain endothelial cell response. The authors nicely combine functional barrier assays with IF for junctional localization in their experimental design, and they delve into potential mechanisms of Cldn5 internalization using markers of endosomal-lysosomal pathway localization. The authors also describe a new type of barrier assay using a streptavidin-coated plate upon which barrier-forming cell cultures can be placted, this could be a very useful alternative or complement to other size-selective barrier assays and presumably could work for other barrier forming cells types, likely epithelial cells.

      Weaknesses:

      (1) There are no measures of bacterial burden in peripheral organs, blood, in the LPM or brain in the TLR4 endothelial cKO mice. Lack of TLR4 in endothelial cells could prevent bacterial 'access' into the LPM and brain, essentially preventing meningitis and leading to a lack of inflammatory responses in the LPM-located cells simply because there is no bacteria present. Bacteremia may also be reduced, as might inflammatory responses in peripheral organs with TLR4-deficient peripheral endothelium. Bacterial counts and inflammatory measures in peripheral organs and blood are important to better understand the mechanism(s) underlying the reduced inflammatory profile in LPM cells and no LPM endothelial breakdown in the Tlr4 endothelial cKO mice. In other words, does deleting TLR4 in EC protect against the development of meningitis by somehow blocking bacteria access to the LPM (this would be supported by low or no CFU counts in infected Tlr4 endothelial cKO) or is it what the authors appear to propose in Figure 1J that TLF4 in EC is the only cell responding to the bacteria to trigger the immune cascade in the LPM? More data is needed to resolve this, as this is a major claim of the paper.

      Thank you for this comment. We agree that it is important to distinguish whether the reduced inflammatory response in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice reflects altered bacterial burden versus altered host sensing. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4CKO mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs of infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice in E. coli burden; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and the time of sacrifice 24 hours later (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid leptomeningeal cells does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (2) The authors look at the underlying cortical response (cerebral vasculature for ICAM and immune cells) but do not use markers that could identify microglia (Iba1), the primary resident immune cell (CD206 is not useful, at this stage, in perivascular macrophages that are extremely sparse in the postnatal brain). This would be important to better study the impact on CNS resident immune cell morphological activation.

      Thank you for this comment. In response, we have analyzed Iba1 staining in the cortex in infected vs. uninfected mice. This is shown in Figure 2 – figure supplement 3. These data demonstrate a several-fold increase in Iba1 immunostaining in infected compared to uninfected cortex, consistent with increased microglial activation in response to infection. There is no statistically significant difference between infected WT and infected Cdh5-CreER; Tlr4CKO mice in Iba1 staining in cortex.

      (3) The authors suggest that Cldn5 junctional localization is selectively disrupted upon bacterial exposure, mediated by TLR4 - they suggest this based on studying PECAM, GLUT1, ZO-1 and B-catenin (all normally junction or cell surface located in cultured Bend3.1) in relationship to Cldn5 localization (normally high) - it is possibly these are also impact by bacteria exposure (maybe through different mechanisms?) - a better measure would be to use the similar cyto/PM measure they do for Cldn5 in Fig. 4D and to evaluate this or to use intensity measurements.

      Thank you for this comment. As the reviewer noted, the analysis of Cldn5 localization with vs. without E. coli exposure and in WT vs. Tlr4KO bEnd.3 cells (shown in Figure 4B and D) – uses Cell Trace to partition the image into cytoplasmic vs. plasma membrane territories. For the analyses in Figure 5, we wanted to compare the localization (and potentially re-localization) behaviors of a variety of subcellular markers with the localization and re-localization of Cldn5 following E. coli exposure. By directly measuring the % overlap of the two immunostains, we get that data. We note that the goal of this analysis is to assess relative co-localization with Cldn5 rather than absolute subcellular partitioning of each marker. While this analysis could have been extended to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories, it is clear by visual inspection of Figure 5A-C that beta-catenin, ZO-1, and PECAM1 remain plasma membrane-associated with E. coli exposure, and GLUT1 goes from the part of the plasma membrane not involved in cell-cell contact without E coli exposure to cytoplasmic with E. coli exposure (as judged by the appearance of a nuclear “shadow” after E. coli exposure). Thus, we do not believe that additional cytoplasmic vs. plasma membrane quantification for these markers would alter the interpretation. The main reason that we did not extend this analysis to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories is because that would introduce the Cell Trace localization as an additional variable.

      (4) The discussion could benefit from delving more into the prior literature on E coli mediated breakdown of junctions in cultured human microvascular brain endothelial cell model and critical host-pathogen interactions of the bacteria with ECs (PMID: 14593586), and how this might involve TLR4.

      Thank you for this comment. Two paragraphs addressing the prior literature have now been added to the discussion.

      (5) It would be important to discuss how their results relate to earlier studies on TLR4-/- and TLR2-/- global knockout mice and protection vs vulnerability to development of meningitis (see PMCID: PMC3524395) - this paper showed that TLR4 global KO mice have increased susceptibility to die from meningitis and have much higher CFU counts in the CNS. In this manuscript and their prior work (Wang et al., 2023), this group shown that both global TLR4-/- mutants and their EC-specific KO have reduced barrier permeability, but we don't have any information about CFU or susceptibility to death from meningitis in their models.

      Thank you for these comments. The model we use – subcutaneous injection of E. coli (a clinical isolate from an infant with meningitis) at postnatal day (P)5 – results in the death of the infected mouse within 2 days (shown in Figure 1 – figure supplement 3 in Wang et al. 2023). Our analyses of infected mice were conducted 24 hours after infection. As noted in the reply to comment #1, in the revised manuscript we present a clinical assessment of WT vs. Cdh5-CreER; Tlr4CKO mice 24 hours after infection based on (1) a quantitative microscopic analysis of E. coli burden in the brain (visualized based on RFP fluorescence in the E. coli used here), (2) quantifying CFUs in blood and (3) mouse weights at P5 and P6, a sensitive indicator of overall health since this is a time when mice are normally gaining weight rapidly (~25% weight gain per day). These data (shown in Figure 2 figure supplement 4) indicate that bacterial burden and disease severity are similar between genotypes in our model. In Wang et al., 2023, we did not conduct a quantitative clinical assessment of WT vs. Tlr4-/- mice following infection, but by visual inspection, infected WT and Tlr4-/- mice appeared to have similar downhill clinical trajectories. We have expanded the Discussion to relate these findings to prior studies of global TLR4 and TLR2 knockout mice, noting that differences in experimental models and the distinction between global versus VECadCreER-specific deletion may account for the differing outcomes reported.

      Comment on the paper listed by the reviewer (PMCID: PMC3524395).

      The cited study demonstrates that global TLR4 deficiency leads to increased bacterial burden and mortality, indicating an essential role for TLR4 in host defense and bacterial clearance. In our study of Cdh5-CreER; Tlr4CKO mice, bacterial burden and disease severity at 24 hours post-infection are similar between WT and Cdh5-CreER; Tlr4CKO mice, indicating that Cdh5-CreER; Tlr4CKO does not alter the clinical course at this time point. This difference is noted in the Discussion section.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the molecular underpinnings of immune responses in the leptomeninges in neonatal bacterial meningitis. Bacterial meningitis is a major disease burden, particularly for neonates, and it has previously been noted that the meningeal immune environment in infants is permissive to opportunistic infection (Kim et al., Sci Immunol, 2023). There is less known about the contribution of the stromal compartment to meningeal immune responses. Seegren et al. interrogate the role of leptomeningeal endothelium in host defence in E. coli infected neonatal mice using mouse genetic tools to delete the LPS receptor Tlr4 from either endothelial cells (using Cdh5-CreER) or macrophages (using LysM-Cre). The authors use snRNAseq, cleared cortical mounts, and in vitro work to define the impact of E. coli infection on leptomeningeal endothelial cells. This study uses a range of innovative techniques to probe the role of the stromal compartment in meningitis.

      Strengths:

      This study makes excellent use of cleared cortical mounts to examine the biology of the leptomeninges, in particular, changes to the endothelium, with unprecedented detail. In combination with high-quality sequencing data provide new insights into the impact of meningitis on the leptomeninges. The data presented by the authors is of very high quality.

      Weaknesses:

      The weaknesses of the study were in terms of interpretation and perhaps study design.

      (1) Most importantly, the authors need to provide additional validation of their conditional knockout models. The authors need to confirm that the Cdh5-CreER does not impact leptomeningeal fibroblasts and to confirm gene deletion in macrophages.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclearlocalized GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite correct. The Results section of the revised manuscript includes an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      (2) The authors could also strengthen the paper by providing data on the impact of these conditional knockout models on the course of meningitis and bacterial burden.

      Thank you for this comment. We agree that these additional analyses strengthen the manuscript. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences in bacterial burden between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) Finally, it is perhaps not surprising that Tlr4 is required for meningitis responses with E. coli. However, it is unclear if these findings can be generalised to other, more common, meningitis infections (streptococcal/pneumococcal).

      At present, it is an open question whether TLR4 plays as a large a role in meningitis caused by other gram-negative bacteria and whether TLR2 plays a similarly large role in meningitis caused by gram-positive bacteria. In the Discussion, the last two sentences under “Limitations of the study” summarize this point: “Finally, the present study focused on E. coli K1, the dominant Gram-negative neonatal pathogen. Future work could assess TLR signaling in response to other bacterial pathogens, such as Group B Streptococcus.”

      (4) There are additional minor issues; for instance, the arachnoid fibroblast 2 population appears to closely resemble dural border cells.

      Thank you for this comment. That is correct, and we have changed the nomenclature to “dural border cells”.

      (5) The cell line model (bEnd.3) is a relatively low-fidelity model of BBB endothelial cells, and this should be acknowledged.

      Thank you for this comment. That is correct. Despite being brain-derived, bEnd.3 cells have lost many BBB-specific attributes. Their responses might best be considered as generic endothelial responses rather than brain-specific endothelial responses. This is now stated in the Results section: “Although they are brain-derived, bEnd.3 cells lack many BBB-specific attributes and, therefore, they likely exhibit generalized endothelial responses rather than brain-specific responses to bacterial exposure.”

      With these caveats, it is difficult to be certain that the endothelium alone is the driver of meningeal immune responses in meningitis, and what the impact of these is.

      We agree with this critique. As noted above, the expression of Cdh5-CreER in essentially all endothelial cells and in a subset of dural border cells and leptomeningeal fibroblasts means that the comparison of TLR4 CKO with Cdh5-CreER vs. Lyz2-Cre is assessing phenotypes driven by TLR4 signaling in endothelial plus a subset of other non-myeloid cells vs. TLR4 signaling in myeloid cells. We have revised the text to reflect this more precise understanding of Cdh5-CreER specificity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Transcriptomic analysis: The analysis and display of the single-nucleus RNA-seq data should be improved. The authors could perform a more granular, unbiased clustering of each cell class in the combined dataset and then compare the proportion of each experimental group (genotype x control/infected) in each cluster. At present, it appears the differentially-expressed genes (DEGs) shown in Figure 1 were identified using the Seurat FindMarkers function with default parameters (Methods). This considers each cell as an independent experimental unit and is therefore not appropriate for a comparison of control versus infected groups (see e.g., PMID 34584091, 35880426. The authors should implement a statistical analysis strategy that considers true biological replicates (mice, as shown in Supplementary File 1).

      We do not fully agree with this critique. We agree that biological replication at the level of individual mice is important for interpreting these data, but within each mouse, the characteristics of individual cells is also of interest, including the degree of heterogeneity, the sample size for a given cell cluster, and the statistical significance of any observed changes in transcript abundance. As requested, we have prepared a new supplemental figure (Figure 1 – figure supplement 5) showing a principal component analysis of the scRNA-seq data for each mouse (one mouse was used for each snRNA-seq dataset) and for each of the principal leptomeningeal cell types. This analysis shows, for example, that the three infected Cdh5-Cre; Tlr4flox/- mice have transcriptomes for each of the six cell clusters that are very similar to the transcriptomes of the two uninfected WT and the two uninfected Cdh5-CreER; Tlr4flox/- mice. Thus, the genotype- and condition-dependent effects are consistent across biological replicates. At the most granular level, Figure 1 – figure supplement 7, which was part of the original submission, shows for the most up- and down-regulated genes (based on adjusted p-value or based on fold-change) in endothelial cells and in myeloid cells how individual transcript abundances change for each mouse and for each genotype/condition.

      (2) The authors use immunohistochemistry to assess claudin-5 "disorganization and redistribution" (Results, pages 7-8 and Figures 3A-B). They state that "Tlr4ECKO mice showed minimal changes in the distribution of Cldn5, implying that cell autonomous endothelial TLR4 signaling regulates tight-junction organization." It is not clear, however, that the quantified parameter (Cldn5+ area relative to total area) would be an accurate readout of claudin-5 organization/distribution (i.e., subcellular localization) as it would also be sensitive to claudin-5 expression, vascular density, and vessel diameter. The authors use a similar assessment of ZO-1 to suggest that changes to claudin-5 are not due to a "generalized disassembly of TJs" and could also use this to argue that the above potential confounds (vascular density, vessel diameter) do not change, but the data in Figure 3 - Figure Supplement 1B, lower panel, show that infection does cause an increase in ZO-1 area relative to total area (P = 0.0004). Thus, the statement in the results "Zonula Occludens-1 (ZO-1) [...] remained unchanged during infection (Figure 3 - figure supplement 1)" is not accurate. The authors should revise this section to ensure their conclusions are aligned with the presented data.

      Thank you for this comment. The reviewer is correct that the Cldn5 area measurement is unable to deconvolve the various factors that might contribute to it (vessel density and diameter, and Cldn5 distribution). This part has been rewritten. “Consistent with prior findings (Wang et al., 2023), both WT and Tlr4<sup>MKO</sup> mice showed an increase in the area occupied by Cldn5 in the leptomeninges following infection, likely referable to both increased vessel diameter and a redistribution of Cldn5 within ECs (Figure 3A-B; Figure 3 – figure supplement 1C).”

      The reviewer is also correct about our initial description of the ZO-1 data. What we meant to write and what the revised manuscript now shows is: “The area occupied by Zonula Occludens-1 (ZO-1), a tight junction scaffold protein, showed a modest but statistically significant increase in WT leptomeningeal vessels but no significant change in Tlr4<sup>VEKO</sup> leptomeningeal vessels during infection (Figure 3 – figure supplement 1A and B).”

      (3) In the Methods, under Mouse Models and E. coli Infection, the authors state, "The Cdh5-CreER line (Monvoisin et al., 2023) was the same line used in Wang et al. (2023)." However, there is no mention of Cdh5-CreER in Wang et al. (2023). Could authors please clarify? Also, because this appears to be an inducible Cre, the authors must include details on the dose and timing of tamoxifen or 4-OHT used in this study.

      Thank you for catching that error. We meant to reference Wang et al (2025), not Wang et al (2023). [Wang et al (2025) is: Wang Y, Rattner A, Li Z, Smallwood PM, Nathans J. (2025) Vascular endothelial-specific loss of TGF-beta signaling as a model for choroidal neovascularization and central nervous system vascular inflammation. Elife 14:RP107018.] This has now been corrected.

      We have now included the details related to 4HT injection in the Methods section “Mouse Models and E. coli Infection”. These are intraperitoneal injection at P2 with 40- 50 µL of 2 mg/ml 4HT.

      (4) The legend for Figure 1A is "Schematic of the leptomeninges", but the figure shows the entire brain-skull interface, including underlying cortex, leptomeninges, dura, and skull.

      Thank you. Corrected.

      (5) Page 7, typo: "In the brain, CD206+ cell were too sparse ..." Should be "cells".

      Thank you. Corrected.

      (6) Page 14, typo: "... could represents a double-edged ..." Should be "represent".

      Thank you. Corrected.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform CFU counts from LPM, dura, brain, peripheral organs (liver) in infected v mock mice from control v TLR4 EC-cKO.

      Thank you for this comment, with which we agree. We have addressed this by quantifying bacterial burden and assessing disease severity in WT and Cdh5CreER; Tlr4floxed mice. Specifically, we performed CFU measurements in blood, monitored mouse weights at P5 and P6, and histologically surveyed the E. coli-RFP signal (i.e., E. coli burden) in brain, liver, and lung. These analyses show that bacterial burden and disease progression are comparable between WT and Cdh5-CreER; Tlr4floxed mice at 24 hours post-infection. These data are presented in Figure 2 – figure supplement 4 and described in the Results.

      (2) Lyz2Cre/+ is used to delete TLR4 from macrophages, but recombination efficiency (in LPM BAMs) is described as only partial, suggesting that TLR4-response in LPM BAMs (and potentially macrophages in the dura) is at least partially intact. It undercuts conclusions that can be made using this line.

      Thank you for this comment. We have conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The Results section text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Also, as noted in the reply to comment 4 below, a direct analysis of Tlr4 recombination efficiency is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (3) Inflammatory responses [qPCR] from peripheral organs and also physiological measures in the pups [weight post-infection, time to moribund or death curves] in control v TLR4 EC-cKO and TLR4 mac-cKO.

      Thank you for this comment. We have not conducted a qPCR analysis of inflammatory gene expression in peripheral organs because (1) the dramatic upregulation of these transcripts in the leptomeninges, (2) the presence of E. coli in blood and peripheral organs, and (3) the clinical assessment (cessation of weight gain) all predict that such an analysis would reveal a large up-regulation of inflammatory gene expression throughout the body. More specifically, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (4) The conditional macrophage line is problematic due to the partial recombination. I question the utility of including this unless they can come up with a way resolve the response of recombined TLR4 macrophages vs ones that are not (could they use the single cell data to pick this a part? Are TLR4-null cells and TLR4 'wt' cells transcriptionally similar in the infected condition, suggesting TLR4 is not doing much in the macs, potentially due to alternate TLRs?). There are good BAM Cre lines that have been described [Lyve1-cre would be good for LPM BAMS, the other is Pf4-cre, see https://pmc.ncbi.nlm.nih.gov/articles/PMC7375817/ - just as an FYI for the future].

      Thank you for this comment. As noted in the reply to point 2 (above), we have conducted a more in-depth analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      We agree that, based on Figure 6 in the cited paper [McKinsey et al (2020) A new genetic strategy for targeting microglia in development and disease eLife 9:e54590], the Pf4-Cre line may be superior to the Lyz2<sup>Cre</sup> line that we used for recombination in leptomeningeal myeloid cells. Unfortunately, we missed this paper in our literature searches, probably because it focuses on a microglial CreER line, P2ry12-CreER, and the Pf4-Cre line is not mentioned in the title or abstract. Our decision to use the Lyz2<sup>Cre</sup> line was based on an extensive comparison among myeloid Cre lines showing that Lyz2<sup>Cre</sup> was the most efficient [Abram CL, Roberge GL, Hu Y, Lowell CA. 2014. Comparative analysis of the efficiency and specificity of myeloid-Cre deleting strains using ROSA-EYFP reporter mice. J Immunol Methods 408:89-100.] However, the Abram et al study did not look at the leptomeninges. Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (5) Figure 1 - Figure Supplement 2 - the authors nicely break down the pathway response [NFKB and TNF] in EC and macs, it would be great to have similar information for the fibroblasts (in the main figure or the supplement). Does their inflammatory response show a similar pattern?

      Thank you for this suggestion. We have now done that analysis and present it in Figure 1 – figure supplement 3. For completeness, we also performed the same type of analyses for JAK-STAT signaling and IFN-gamma response and these are shown in Figure 1 – figure supplement 6. The principal conclusion is that across all major leptomeningeal cell types, the Cdh5-CreER; Tlr4floxed samples (i.e., Tlr4 KO’d in non-myeloid cells) show much reduced transcriptome changes with infection.

      (6) What is ASC and Cd206 quantification measuring, and how does this relate to 'activation' - is this the intensity of signal or a morphological change? What is the precedence for using ASC (citations)? In their prior work, they showed no change in CD206 number, so a significant increase upon infection here, it's confusing exactly what is being studied. Also, loss of Lyve1 is a well-accepted measure of activation that they have previously used, adding that it could be helpful. This is not a major issue since they have robust data that the macrophages are not transcriptionally activated. Clarification of what exactly is being measured would be sufficient (in the text).

      CD206 immunostaining, which reveals myeloid cell morphology, shows that, with E. coli infection, myeloid cells convert from a more compact morphology to a more expanded morphology. This is now explained more fully in the Results section.

      Regarding ASC, changes in the state of ASC aggregation and ASC subcellular localization have been used by others to monitor immune cell responses to inflammatory signals (Sester et al., 2016; Franklin et al., 2018). While this change in subcellular localization may explain part of the increase in immunostained area in myeloid cells in the infected mice (Figure 2D), the increase in the area of ASC immunostaining largely reflects a shift of myeloid cells from a compact to a more extended morphology. This is now explained more fully in the Results section. We have also added two references (Sester et al., 2016; Franklin et al., 2018) that described how ASC distribution changes with inflammation.

      Regarding LYVE1, we observe a decrease in LYVE1 transcript abundance in myeloid cells with infection, as predicted. Given the large amount of other data that document myeloid activation with infection, we have elected not to include this.

      (7) The authors suggest the internalization of Cldn5 is not due to NFKB downstream signaling that includes transcriptional mechanisms because it happens as early as 1 hour, prior to NFKB localization to the nucleus. However, a lot of their experiments, including on endosomal-lysosomal protein co-localization are done at 4 hours, when their RNAseq data show robust NFKB-mediated gene upregulation and (though not tested) potentially protein production of factors that can act back on the cells, including to impact endo-lysosomal processing. Without studies at earlier timepoints post-bacteria exposure, separating these two mechanisms is difficult.

      Thank you for this comment. We have explored this question by looking at Cldn5 internalization in bEnd.3 cells at 1 hour after E. coli exposure, and the data clearly show that internalization occurs within 1 hour. Additionally, we have conducted this experiment in the presence of 1 uM ACHP, an IKK inhibitor that blocks NF-кB migration to the nucleus. ACHP treatment shows no effect on the rapid internalization of Cldn5, implying a mechanism independent of NF-кB control of gene expression. These data are shown in a new figure (Figure 6) in the revised manuscript.

      (8) Figure 2 - CD206 are quite sparse however, Iba1 would work well to look at microglial activation.

      Thank you for this suggestion, which we have followed. To assess microglial activation, we have immunostained for Iba1 and quantified the data. These are now included in Figure 2 – figure supplement 3. The data show that there is an increase in Iba1 immunostaining following E. coli infection in both WT and Cdh5-CreER; Tlr4floxed mice, with more in the former than the latter, but the difference is not statistically significant.

      (9) Suggest performing the LAMP+ co-localization experiment at <1hr, prior to NFKB nuclear localization and transcriptional changes. This would better support it, this is (or is not) independent of the NFKB. Could also test this with an NFKB inhibitor, do they still see the CLDN5 internalization when NFKB is blocked?

      Thank you for these suggestions. We have done both of these analyses, and the results are presented in Figure 6. The results show that (1) Cldn5 is internalized within 1 hour and (2) its internalization is independent of NF-кB signaling inhibition by 1 uM ACHP. Since ACHP treatment shows no effect on the rapid internalization of Cldn5, that implies a mechanism independent of NF-кB control for gene expression.

      Reviewer #3 (Recommendations for the authors):

      Major points

      (1) The most important caveat is that the Cdh5-CreER model is known to recombine in leptomeningeal fibroblasts (10.1038/s41586-023-06993-7, 10.1101/2025.05.13.653681), and Cdh5 expression in these populations is now well described (10.1038/s41467-02341580-4, 10.1016/j.neuron.2023.09.002). Although the authors did not observe recombination in their reporter (details of the tamoxifen injection protocol should be provided), it is imperative to validate the specificity of their model to Tlr4 in endothelial cells, leveraging their sequencing data and providing additional IHC or ISH to confirm this. Alternatively, Tlr4 could be deleted in a more specific model, e.g., the Pdgfb-iCreERT2 or Slco1c1-CreERT2. It is also important to do the same with the LysM model, to confirm that the lack of impact of macrophage Tlr4 is not due to failure to delete the gene. This is again important to the interpretation of the study, since the authors propose that the endothelium, specifically, is the driver of the meningitis response.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclear-localised GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite right. The Results section of the revised manuscript has an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      We have also conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup>-directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text in the Results section now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. The phenotype of Cdh5-CreER; Tlr4floxed mice – a dramatically reduced infection-associated transcriptional response – argues that the floxed Tlr4 target was recombined at appreciable efficiency in those mice (Figure 1D and 1E). For Lyz2<sup>Cre</sup>; Tlr4floxed mice the principal phenotype is an up-regulation of infection-associated transcripts in a subset of dural border cells in the absence of infection; the transcriptional response to infection was largely unaffected in all leptomeningeal cell types (Figure 1D and 1E). We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4VEKO and Tlr4MKO mice.”

      (2) The authors did not examine the consequences of Tlr4 cKO on the course of meningitis or bacterial burden. Knowing the impact of this would strengthen the paper and allow us to determine if the endothelial responses are helpful or harmful in meningitis progression.

      For the revised manuscript, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) TLR4 is a known receptor for LPS. It is unsurprising (especially in the in vitro experiments) that Tlr4 knockout reduces NF-kB signalling and other downstream changes to endothelial cells. Furthermore, it is uncertain if the infection was left to continue, similar changes to the endothelium would nonetheless occur through other mediators such as IL1B and TNFa.

      We agree that it makes logical sense that Tlr4 KO decreases NF-кB signaling. The interesting next question is: what are the mechanistic underpinnings of the responses that are downstream of TLR4 and NF-кB? The cell culture experiments with WT vs. Tlr4KO bEnd.3 cells identify one set of cell biological responses related to Cldn5 and junctional integrity, and the NF-кB inhibition experiment (Figure 6) implies that rapid internalization of Cldn5 occurs in the absence of NF-кB mediated transcriptional changes. Regarding the possibility that other mediators such as IL1B or TNFα might, at least partially, make up for the lack of TLR4 signaling later in the infection, that is an open question at present.

      (3) The arachnoid fibroblast 2 cluster should be renamed to dural border cells based on their high expression of Slc4a10, Adamtsl3, Tmeff2, etc which are all highly enriched in dural border cells. I suspect this cluster is also highly enriched for Slc47a1, probably the most specific marker for these cells (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002).

      Thank you for this comment. The reviewer is correct. These are dural border cells and they express Slc47a1, as seen in a new supplemental Figure 1 – figure supplement 4, which shows UMAP plots for many leptomeningeal cell type-specific genes. We have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (4) It would be helpful to provide higher resolution images of Cldn5 in the leptomeningeal mounts. At the current resolution, it is difficult to tell if there is a similar internalisation/disruption phenotype to what is observed in vitro. Notably, this finding is similar to another recent publication on Cldn5 recycling (in the context of stroke) (10.1186/s40478-025-02125-6).

      Higher resolution images of Cldn5 in leptomeningeal vessels without or with E. coli infection are now shown in Figure 3 - figure supplement 1C. There is a visual impression of greater area occupied by Cldn5, which is confirmed by quantification (Figure 3A and B). This effect appears to be due to both an average increase in vessel diameter and a redistribution of some of the Cldn5 away from plasma membrane junctions. Thank you for pointing out the interesting and relevant Cottarelli et al (2025) paper, which we had not read. This is now referenced.

      Minor points

      (1) Typo: prominant should be spelled prominent.

      Thank you for catching that one. It is now corrected.

      (2) Strictly speaking, the arachnoid layer is not epithelial (despite Cdh1 expression). They are fibroblasts that acquire barrier-forming properties.

      Thank you for that comment. That appears to be the consensus view, and we will go along with it.

      (3) Notably, LyzM Cre will also recombine in other myeloid populations, so I wouldn't describe it as a macrophage.

      Thank you for this comment. We agree, and we have therefore changed the text and figure labels from “macrophage” to “myeloid”.

      (4) It is interesting and notable that ICAM1 expression is observed in nonendothelial populations, in the IHC, too, perhaps.

      We agree. ICAM1 may be a broader marker/mediator of inflammation than is generally recognized.

      (5) In F1B, your labels on the right image to the arachnoid barrier and pial surface are presumably meant to refer to the image on the left with DPP4 and laminin labelling? The subarachnoid should be between the laminin and DPP4 layers (although it will be collapsed in your preparations).

      Thank you for catching this error. The vertical bars were sized erroneously, and the labels were also placed erroneously. These have now been corrected.

      (6) I would reference the papers that defined leptomeningeal cell type markers (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002) when you define your cell types.

      Thank you. We have done that, and we have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (7) I would change references to the subarachnoid space in your figures to the leptomeninges (which include the SAS, but extend either side of it).

      Thank you. The labels have been changed to “leptomeninges”.

      (8) In Figure 2 - Supplement 1A, it looks like the populations are mislabelled.

      Thank you. This has been corrected to be consistent with the assignments in Figure 1B

    1. eLife Assessment

      This important study addresses how the molecular identity of a single neuron specifies its hard-wired synaptic connectivity, using repeated single-cell RNA sequencing of identified Drosophila sensory neurons together with functional perturbation of candidate cell-surface molecules. It demonstrates remarkably low transcriptomic variability across animals for the same identified neuron, defines a tractable set of differentially expressed cell-surface molecules that distinguish mechanosensory from chemosensory neurons, and links several of these molecules to axonal targeting and circuit function. The evidence is solid, with the single-neuron transcriptomic datasets and Dscam isoform repertoires offering a lasting resource for the field, though clearer articulation of the experimental logic, additional controls in the RNAi screen, and a more complete characterization of the neuronal re-wiring would further strengthen the central claims.

    2. Reviewer #1 (Public review):

      Summary:

      The authors sequence the transcriptome of three sensory neurons from D. melanogaster to study the cell-cell and animal-animal variability in these cells, with a focus on cell adhesion molecules. The work reports useful cell-specific transcriptomics datasets that will be of great interest to those studying cell types, transcriptomes, neuronal development, and cell surface proteomes. The authors also report large numbers of knockdown data (gene-by-gene or in combinations) and report neuronal wiring and behavioral phenotypes. The manuscript is highly descriptive of the system studied - in a good way, but often over-speculates in rationale or conclusions.

      Strengths:

      The manuscript is data-rich. The single-cell transcriptomics datasets, not trivial to collect, are a major strength of the work and will prove useful to the field. Also, the biased expression of Dscam is interesting, even though the authors cannot pursue the mechanism or a function for this.

      Weaknesses:

      The study lacks depth (i.e., mechanism) in explaining observations.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, dos Santos et al seek to identify cell-specific programs that drive neuronal wiring patterns. They focus on two chemosensory and mechanosensory neurons in the Drosophila nervous system, as they both display stereotyped connectivity in the ventral nerve cord. Single-neuron RNA sequencing identified cell surface molecules that distinguish the sensory neurons and may instruct their respective wiring patterns. They functionally test several of these candidates and observe miswiring phenotypes upon knockdown experiments. Additionally, they attempt to miswire the chemosensory neurons. Overall, this manuscript addresses an important question about how neurons identify appropriate synaptic partners through precise cell surface molecular codes. However, there are significant deficiencies in the experimental logic and rigor, and the manuscript can be very difficult to digest.

      Strengths:

      The use of two sensory neurons with stereotyped connectivity is a significant strength, as this enables the authors to identify genes that are required for wiring. Additionally, analyzing the transcriptomes of single neurons repeatedly could potentially be a robust approach to identifying cell-specific cell-surface molecules that drive wiring.

      Weaknesses:

      (1) The authors perform RNAseq for single identifiable neurons, as opposed to neuronal subclasses, which has been reported before. It would be beneficial to elaborate on the significance of using single neurons for answering the scientific question. This is briefly mentioned toward the end of one of the results subsections: "Repeated RNA sequencing of an identifiable neuron seeks to address the fundamental nature of variability in connectomics, axonal branching, and cellular identity." But this should be in the Introduction.

      (2) The authors chose the P14 pupal stage for one of the analyses. It is not clear why this specific stage is chosen. Does pSc and aPa connectivity occur at this stage?

      (3) This reviewer is confused as to why looking at differentially expressed CSMs between pupal and adult stages of two different neurons is useful. This does not seem like an appropriate comparison. This data might be better in the supplemental material, especially given the lack of precise age synchronization across pupal samples (as reported).

      (4) It is very difficult to follow the logic because the manuscript seems to jump around between different results and lacks a compelling through line.

      (5) "Single cell sequencing of the same neuron reveals transcriptome precision": What are the controls here? An aPa neuron is shown in Figure 3 as an example of a different neuronal subtype, but were other factors (e.g., lack of Repo expression) checked to ensure that samples were not contaminated?

      (6) "However, whether any of these exon 6 or 9 splicing specificities are biologically significant can only be determined using exon 6 and 9 isoform-specific RNAi." The authors could alternatively use CRISPR techniques to target specific isoforms that they hypothesize might be important for neural wiring, enabling them to assess isoform-specific wiring defects.

      (7) In the section "The set of cell surface receptors required to wire up the pSc mechanosensory neuron": Several previous subsections of the Results use RNAseq to identify molecules expressed in pSc neurons across different stages. It's unclear why the authors did not start with the identified list of candidate cell surface receptors identified in their RNAseq experiments.

      a. Were any of the genes screened the same as those identified by the authors as differentially expressed in pSc mechanosensory neurons, either across developmental stage (pupa vs. adult) or across neuronal subtype (pSc vs. Gr59d)? If so, it would be helpful to state this here. (They do mention later on that five CSMs identified were more highly expressed in pSc than aPa. However, changes in expression across developmental stages within the pSc neuron would still be helpful to comment on, especially since the authors identified greater transcriptomic differences across developmental stages than they did between different neuronal subtypes.)

      b. The 39 genes not expressed in pSc neurons served as their negative control, but the average axonal targeting grade was 2.3 (between moderate and severe). This calls into question the use of this method as an appropriate measure of whether a gene expressed by pSc neurons is truly required for proper axon targeting; there seems to be a strong probability of significant off-target effects. Performing a global knockdown and cell-specific rescue could potentially complement these experiments and serve as a stronger indicator of candidate receptors' roles in pSc-specific axon targeting.

      (8) It seems as though the purpose of the experiments described in the last results subsection ("Re-wiring the Gr59 chemosensory neuron") is to redirect the Gr59d neuron toward the pSc neuron's axonal targeting phenotype. However, the authors do not state whether they were able to do so effectively (i.e., whether or not there were significant differences between the rewired Gr59d neuron and the pSc neuron). This leaves the story unfinished.

      (9) At the end of the discussion, the authors state that "...if a Gr59d chemosensory neuron is functionally rewired to a pSc mechanosensory circuit, activation of the Gr59d neuron using a bitter tastant molecule should elicit a grooming (mechanosensory) response...". The authors should attempt this experiment, especially given that they have developed the PXGS technique.

    4. Author response:

      We are pleased that the reviewers found the repeated single-neuron sequencing and the finding of less than 1% transcriptomic variability to be original and striking, valued the single-neuron Dscam isoform repertoires and the scale of the functional screen, and judged the evidence solid to compelling. We provide below our provisional response and an outline of the revisions we plan.

      Overall plan: We intend to submit a revised version that addresses the public reviews and the recommendations to the authors. Because our conclusions rest on data already in the manuscript, the revisions are clarifications, added analysis of existing data, tempered language, and improved figures, rather than new experiments. Given the focused nature of these revisions, we would be happy for the editors to assess the revised version without re-involving the reviewers.

      One factual note for the Assessment and public reviews: The morphological RNAi screen comprised 213 cell-surface receptor genes; the figure of “140 genes” in one public review is the number that produced strong-to-severe phenotypes (Grade 3–5 at >40% penetrance), not the number screened. We will make this unambiguous in the revised text.

      Main changes in the revision:

      (1) We will explain that the RNAi screen was performed blind and independent of the RNA sequencing experiments. This was intentional, so that functional perturbation and transcriptomic identity would serve as independent lines of evidence, but could be compared with each other.

      (2) We will revise the Methods and Results to clarify how the morphological RNAi screen and behavioral subset should be interpreted, including conservative treatment of the negative-control distribution and mild-to-moderate phenotypes.

      (3) We will soften language that overstated certainty. Differentially expressed molecules are now described as prioritized candidates and convergent evidence, not definitive determinants.

      (4) We will reframe Gr59d/PXGS experiments as morphological rewiring and ectopic branching, and no longer imply a pSc-like conversion or functional rewiring.

      (5) We will add a limitations paragraph addressing RNAi off-target/background concerns, the absence of direct aPa functional testing, and the need for future mechanistic validation.

      (6) We will disclose or remove any figure panels that overlap with the companion PXGS manuscript and revise legends/labels to make it more clear.

      We hope these revisions make the logic of the study clearer and align the strength of the claims with the evidence.

      Below is our more detailed (provisional) response (not sure if this is required at this stage):

      Response to the eLife Assessment:

      Clearer articulation of the experimental logic. The Assessment’s central request (Reviewer 2) concerns the relationship between the transcriptomic experiments and the functional screen. The two were performed independently on purpose; the RNAi screen was assembled from a comprehensive survey of the literature rather than from the results of our differentially expressed genes from single-cell RNA sequencing. Thus, the RNAi screen was performed and graded blind in parallel with the single cell sequencing, with the gene identities unmasked only after both were complete. This was intentional, so that the sequencing (i.e., which molecules differ between neurons) and the screen (which molecules are functionally required) would provide mutually unbiased corroborating evidence rather than self-referential support/circular reasoning. We will state this more explicitly in the Introduction, in the Results where the screen is introduced, and in the Methods.

      Additional controls in the RNAi screen. We will treat the 39 genes that were identified in our single-cell RNA sequencing to be not expressed in the pSc neuron as a randomized negative control set in our RNAi experiments. We will state more explicitly the empirical RNAi false positive rate for a miswiring phenotype is 6/39 = 15%, likely due to RNAi off-target effects. We will also more clearly state that our claims about cell surface receptor functions are restricted to strong-to-severe phenotypes at high penetrance reproduced by at least two independent RNAi lines and corroborated independently (differential expression and/or single neuron qPCR).

      A more complete characterization of the re-wiring. We will state more clearly that mis-expressing the pSc-enriched cell surface receptors within Gr59d neurons partially shifts the arbour toward a pSc-like pattern (e.g., increased ectopic branching), and does not reproduce the full anatomical wiring, and that functional/behavioral re-wiring was not tested.

      Response to Reviewer 1:

      Reviewer 1 found the work valuable and data-rich, and the Dscam expression bias interesting. They noted over-confident language and asked how rigorously the differentially expressed genes were identified.

      Over-confident language. We will rewrite the two flagged sentences. The claim that the ~10 differentially expressed molecules are “likely the most important” will become a correlational statement, while also noting the lack of an aPa-specific Gal4 driver for direct testing. Our sentence that, “Our RNA sequencing data is biologically inadequate without a functional characterization of each molecule within the specific neuron” will be replaced with a clearer statement that gene expression data can nominate candidates, and functional perturbation of each gene is required to demonstrate necessity and sufficiency (i.e., biological function); which is exactly why we paired the RNA sequencing with an independent RNAi screen.

      Rigor of the differential-expression calls. We will more clearly state the statistical criteria in the Results (absolute log2 fold change ≥ 2 and Benjamini–Hochberg-adjusted p < 0.05). We will also note the small replicate numbers for the pooled pSc versus aPa comparisons, and emphasize that the central gene calls are independently supported by the blind RNAi screen and, for five genes, by single neuron qPCR. The full statistical workflow is in the Methods.

      Response to Reviewer 2:

      Reviewer 2 considered the findings potentially important but raised concerns about the experimental logic, the rigour of the screen, the completeness of the re-wiring, figure quality, and overlap with our PXGS companion paper. We will address each.

      Experimental logic. Beyond the design of our independent, blinded RNAi screen described above, we will add to the Introduction the rationale for sequencing single identified neurons (rather than subclasses) along with the two-pronged strategy, and add a summary paragraph at the start of the Discussion.

      Developmental stage choices. We will clarify our justification for the P14 pupal stage (the period when the mechanosensory neuron is actively elaborating its arbour while also enabling dissection). We will also clarify the rationale and caveats for comparing the pupal pSc neuron with the adult Gr59d neuron (i.e., the wiring occurs at the pupal stage, but the pupal Gr59d neurons could not be isolated at sufficient quality; the pSc pupal samples are less age-synchronized, so we simply used the comparison to identify the genes shared with the adult comparison).

      Transcriptome precision controls. We will state that the ten single pSc neurons passed the same quality controls for neuronal markers (elav, nSyb) and glial markers (Repo, moody < 20 reads) as all single-neuron libraries, which argues against any contamination by the attendant glial cell, and the aPa transcriptome is used as a different identity comparison.

      Off-target rate. As stated above, we will add the false positive rate for RNAi and restrict our confidence claims to those genes/cell surface receptors with multiple lines of evidence (e.g., strong phenotype, multiple RNAi lines, RNA sequencing, etc).

      Rewiring completeness and the behavioral prediction. As stated above, we will clarify that true re-wiring of the Gr59d neuron requires a future experiment, where a bitter tastant stimulus would elicit a grooming response.

      Response to Reviewer 3:

      We thank Reviewer 3 for judging our work to be fundamental in significance and the evidence compelling, with no major criticisms. Our clarifications above will further reinforce our hypothesis that the differential expression of specific cell surface receptors “do, in fact, control synaptic patterns,” which the reviewer highlighted.

      We are grateful for the reviewers’ time and for eLife’s model. We believe the planned revisions substantially clarify the experimental logic and tighten the claims, and we look forward to submitting the revised version.

    1. eLife Assessment

      This study addresses an important question in liver biology: how zonal hepatocytes balance survival and proliferation following injury? The authors propose that a mid-zone Atf4-Chop axis to Btg2 program temporarily suppresses proliferation to promote survival after a variety of chemical and surgical liver injury models. The authors provide evidence that some zones mount tailored stress responses, which ultimately promote regeneration and liver healing; however, the "mid-zone" changes with different injury models, making it difficult to conclude that the ATF4-CHOP response is specific to this zone in all injury contexts. In addition, it is possible that Atf4 and Btg2 overexpression could lead to Cyp2e1 suppression, which could reduce the extent of injury after CCl4 or APAP. To some extent, these points make the strength of the evidence incomplete, but do not entirely detract from the significance of the study, which is underscored by the helpful observation that there are zone-specific stress responses that mediate liver regeneration and survival.

    2. Reviewer #1 (Public review):

      Summary:

      The authors present evidence that during acetaminophen (APAP)-induced liver injury, mid-zone hepatocytes activate an integrated stress response (ISR) program via Atf4 and Chop, leading to induction of Btg2. This program suppresses proliferation in the early phase of injury, prioritizing hepatocyte survival before regeneration begins. The study uses spatial transcriptomics, immunohistochemistry, CUT&RUN, and AAV overexpression to support this model.

      Strengths:

      (1) Innovative use of spatial transcriptomics to capture zonal differences in hepatocyte stress responses.

      (2) Identification of a mid-zone specific ISR signature and candidate downstream regulator Btg2

      (3) Functional experiments with Atf4-Chop-Btg2 modulation provide causal evidence linking ISR activation to proliferation inhibition.

      (4) Conceptually significant model that hepatocytes actively balance survival and regeneration dynamically in a zone-specific manner.

      (5) Multiple models validation of the finding

      (6) The functional link of such zone2 phenotype is added.

    3. Reviewer #2 (Public review):

      The manuscript reports protection of midlobular hepatocytes from APAP toxicity by activation of Atf4-CHOP (Ddit3)-mediated cell cycle arrest and stress response. The authors acknowledge that their finding is unexpected because CHOP typically induces cell death. Therefore, they functionally validate several aspects of the proposed Atf4-CHOP mechanism. Along these lines, the mitigation of APAP toxicity by AAV expression of Atf4 or Btg2, the latter identified as CHOP effector, is impressive. Whether Atf4 indeed acts through CHOP and whether midlobular hepatocytes are protected because of cell cycle arrest is less clear. These and other criticisms are described in the following.

      Major points:

      (1) Starting with the basics, one wonders why midlobular hepatocytes manage to mount a defensive response to APAP, but PC hepatocytes don't. Is this because midlobular hepatocytes express the relevant Cyps (2e1 but also 1a2 and 3a11) at lower levels, which mitigates toxicity and buys them time? This would be supported by F2A but not by F3B, at least not for the most important Cyp2e1. A moderate difference is shown for Cyp1a2 expression in F3D but is that enough to explain the different fates? Or are additional post-transcriptional effects on these Cyps at work? The difference in baseline Cyp2e1 expression between F2A and F3B remains unexplained after revision.

      (2) The evidence presented in support of cell cycle arrest of midlobular hepatocytes is not fully convincing: there is no overt difference in S and G2/M gene scores in F2F; the marker genes used for S phase and G1 to S progression in F2G are unusual. Along these lines, one wonders if spatial transcriptomics confirmed the Ki67 immunostaining results in F1 also for specific zones, not only overall as shown in F2E? In contrast to the revised discussion, the abstract does not reflect that limited evidence for a cell cycle arrest in pericentral hepatocytes was found.

      (3) The authors conclude in line 364 that halting of proliferation by Btg2 favors survival, which raises the question of whether Btg2 knockout causes death in midlobular hepatocytes in F6K. Data addressing this question, that is, localization and extent of tissue necrosis and ALT levels after APAP, are missing. The efficiency of knockout of Btg2 is also not given. Additional Btg2 knockout data support its proposed role in the revised manuscript.

      (4) Related to the previous question, the BTG2 immunostaining in F6F is not convincing when compared to F6D. One also wonders if it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3? The BTG2 immunostaining remains weak, not only in in F6F but now also in F6D of the revised manuscript, which together with lack of high-resolution immunostaining of AAV-Ddit3-induced BTG2 in the absence of APAP results in limited support for the conclusion that APAP promotes nuclear localization of BTG2.

      (5) Related to the previous question, the proposed Atf4-Ddit3 axis is challenged by the lack of midlobular induction of Atf4 in the APAP scRNA-seq data published by another group presented in S4F and G. Further analysis of AAV-Atf4 samples generated for F5 could address if it is really Atf4 that acts on Ddit3 in APAP toxicity. The extended list of transcription factors (from 30 to 50) includes Atf4 but direct evidence for an interaction with Ddit3 is missing from the revised manuscript.

      (6) Related to the previous question, the ATF4 immunostaining in F5A doesn't look convincing, with many brown pigments appearing to be outside of the nucleus. The ATF4 immunostaining after APAP challenge remains weak.

      (7) It is not ruled out that AAV expression of Atf4 or Btg2 reduces hepatocyte sensitivity to APAP by affecting expression of the Cyps needed for activation. In other words, does AAV-Atf4 or AAV-Btg2 change the expression of any of the Cyps relevant to APAP in the 3 weeks before APAP application (F5B)? S5A of the revised manuscript rules out loss of Cyp2e1 expression as a confounding factor.

      (8) It is laudable that the authors tried to extend their findings to human by using snRNA-seq data from a published study (line 391) but it is unclear why they didn't analyze all 10 patients in that study but instead focused on 2 and stated that this small sample number prevented drawing definitive conclusions and could therefore only be mentioned in the discussion. The revised manuscript continues to focus on rare spatial transcriptomics analyses of patients with APAP toxicity although more snRNA-seq analyses of such patients are available which should also allow for analysis of hepatocyte zonation.

      Comments on revised version.

      After revision, the proposed role of Btg2 is substantiated but it remains unclear why midlobular hepatocytes don't proliferate after APAP challenge and whether the observed protective effects are indeed mediated by Atf4 acting directly through CHOP.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Zonation definition under injury has been shown to be sustained broadly, but is not sufficiently validated and quantified, especially considering the resolution of the 10x Visium system and the potential variation of outcomes based on how to define zones.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g.,Alb, Mup20, Cyp2f2,Pck1,Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2,Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g.,Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) The model is built entirely in APAP injury, which specifically targets pericentral hepatocytes. It remains unclear whether the proposed mechanism applies to other liver injuries (e.g., partial hepatectomy, CCl4).

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Baseline proliferation appears higher than expected in homeostasis (Figure 1B), and fold change analysis (not absolute counts) may be needed to assess zonal proliferation suppression (Figure 1D).

      We thank the reviewer for this insightful comment. The baseline proliferation rates observed in our study are consistent with previously reported zonal distributions (PMID: 33632817; PMID: 33632818), with approximately 70% of proliferating hepatocytes located in zone 2, 20% in zone 3, and 10% in zone 1 under homeostatic conditions. To further address the reviewer’s concern, we performed a fold-change analysis of Ki-67<sup>+</sup>hepatocytes across different zones. This analysis revealed that only the mid (zone 2) and pericentral regions exhibited significant changes, whereas no statistically significant differences were observed in the other zones (as shown in Author response image 1). Importantly, when considered together with the absolute cell counts, these results indicate that the apparent suppression of proliferation is most pronounced in the mid zone, likely due to its relatively higher baseline proliferation under homeostatic conditions. In contrast, this effect is less evident in the fold-change analysis, as zones with low baseline proliferation show limited dynamic range for detecting relative changes.

      Author response image 1.

      Fold changes of Ki67-positive cells across liver zones (PC, Mid, PP) at 0, 3, 6, 12 and 24 h post-APAP. Fold change was the number of Ki67-positive cells in the three regions at each time point after APAP treatment divided by the number of positive cells in each region at 0 hour post-APAP. (a) denotes significance between PC and Mid regions, (b) denotes significance between PC and PP regions, and (c) denotes significance between Mid and PP regions.

      (4) AAV-based overexpression raises potential confounds (altered CYP activity before injury) and shows incomplete penetrance that is not quantified (Figure 5 - Figure 6).

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (5) The functional link between proliferation suppression and improved survival is inferred, but direct survival /injury readouts are limited.

      We thank the reviewer for this insightful comment. To more directly evaluate the functional link between proliferation control and liver injury, we manipulated Btg2, a downstream effector of the Atf4–Chop axis and a known inhibitor of cell proliferation. Knockdown of Btg2 using AAV8–CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial knockdown efficiency, we observed a clear exacerbation of liver injury, as evidenced by an approximately 2-fold increase in serum ALT levels and a ~1.5-fold expansion of necrotic areas. In parallel, hepatocyte proliferation was significantly increased (~1.8-fold increase in Ki67⁺ hepatocytes) compared to control mice (Revised Figure 6K–N). Conversely, Btg2 overexpression produced the opposite phenotype, markedly attenuating liver injury while suppressing hepatocyte proliferation (Revised Figure 6G–J). Together, these gain- and loss-of-function data provide direct evidence linking proliferation control to injury severity, thereby supporting a causal relationship between suppressed proliferation and improved liver outcomes (Revised manuscript, page 14, lines 399–406).

      Reviewer #2 (Public Review):

      (1) Starting with the basics, one wonders why midlobular hepatocytes manage to mount a defensive response to APAP but pericentral hepatocytes don't. Is this because midlobular hepatocytes express the relevant Cyps (2e1, but also 1a2 and 3a11) at lower levels, which mitigates toxicity and buys them time? This would be supported by F2A but not by F3B, at least not for the most important Cyp2e1. A moderate difference is shown for Cyp1a2 expression in F3D, but is that enough to explain the different fates? Or are additional post-transcriptional effects on these Cyps at work?

      We thank the reviewer for this important question. We fully agree that the differential susceptibility between mid‑zone and pericentral (PC) hepatocytes is likely rooted in the zonal gradient of cytochrome P450 expression. Our spatial transcriptomics data (Revised Figure 2A) show that mid‑zone hepatocytes express Cyp2e1, Cyp1a2, and Cyp3a11 at levels intermediate between PC and periportal (PP) zones. This intermediate expression may generate sufficient NAPQI to activate stress signaling but not so much as to cause immediate mitochondrial collapse, thus “buying time” for adaptive responses. We also appreciate the reviewer’s observation that Cyp2e1 mRNA levels remain highest in the PC zone even after APAP (Revised Figure 3B). However, mRNA abundance does not necessarily reflect functional protein level. In the PC zone, massive necrosis rapidly compromises cellular integrity; as shown in Revised Figure 3D, Cyp1a2 protein declines sharply around the central vein, and we observed similar degradation for Cyp2e1 (data not shown). Consequently, despite sustained Cyp2e1 transcripts, PC hepatocytes are unable to mount an effective stress response because they are already undergoing cell death. By contrast, mid‑zone hepatocytes retain sufficient metabolic capacity to activate the Atf4‑Chop axis while preserving cellular function.

      (2) The evidence presented in support of cell cycle arrest of midlobular hepatocytes is not fully convincing: there is no overt difference in S and G2/M gene scores in F2F; the marker genes used for S phase and G1 to S progression in F2G are unusual. Along these lines, one wonders if spatial transcriptomics confirmed the Ki67 immunostaining results in F1 also for specific zones, not only overall, as shown in F2E?

      We thank the reviewer for these important observations. We agree that the current spatial transcriptomics (ST) data alone do not provide sufficiently strong support for this conclusion. The limited sensitivity of ST for detecting rare proliferative events further constrains its utility in this context. At baseline, only ~1% of ST spots are Ki67-positive (Revised Figure S1I), and this fraction becomes even lower during the early phase following APAP injury. As a result, there are insufficient Ki67+ spots to robustly assess zonal distribution using ST, which precludes a reliable spatial validation of proliferation patterns at this resolution. For this reason, our primary evidence for zonal proliferation dynamics relies on Ki67 immunohistochemistry (Revised Figure 1), which provides single-cell resolution and higher sensitivity. These data show a marked reduction in Ki67+ hepatocytes specifically in the midlobular zone at 3-6 hours post-APAP, supporting a transient suppression of proliferation in this region. In addition, we agree that the transcriptional evidence for cell cycle arrest was not strong the S and G2/M scores showed no overt difference, and the gene sets used were suboptimal. We have therefore moved these analyses to the supplement and toned down the claims. We have also clarified this limitation in the manuscript (Revised manuscript, page 18, line 518-524)

      (3) The authors conclude in line 364 that halting of proliferation by Btg2 favors survival, which raises the question of whether Btg2 knockout causes death in midlobular hepatocytes in F6K. Data addressing this question, that is, the localization and extent of tissue necrosis and ALT levels after APAP, are missing. The efficiency of the knockout of Btg2 is also not given.

      We thank the reviewer for this insightful comment. We have included the missing data. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup>hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N, Revised manuscript, page 14, line 399-406).

      (4) Related to the previous question, the BTG2 immunostaining in F6F is not convincing when compared to F6D. One also wonders if it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3?

      We thank the reviewer for this insightful comment. We have included an inset of the original image to better show BTG2 staining in revised Figure 6F. During our study, we tested BTG2 expression in mice transduced with AAV‑TBG‑EGFP or AAV‑TBG‑BTG2 for three weeks without APAP challenge. We observed that BTG2 in these non‑injured livers was predominantly cytoplasmic (Author response image 2), contrasting with the nuclear localization seen after APAP treatment (Figure 6F). Regarding whether it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3, we think Ddit3 promotes BTG2 expression (as shown in revised Figure F6F), but APAP is necessary for its nuclear translocation.

      Author response image 2.

      Immunohistochemical detection of Btg2 in liver tissue from mice transduced with AAV-TBG-EGFP or AAV-TBG-Btg2 for 3 weeks without APAP treatment.

      (5) Related to the previous question, the proposed Atf4-Ddit3 axis is challenged by the lack of midlobular induction of Atf4 in the APAP scRNA-seq data published by another group, presented in S4F and G. Further analysis of AAV-Atf4 samples generated for F5 could address whether it is really Atf4 that acts on Ddit3 in APAP toxicity.

      We thank the reviewer for this insightful comment. We agree that Atf4 was not among the top 30 active transcription factors in our initial analysis; however, when we extended the list to the top 50, Atf4 was included. We have therefore updated Revised Figures S4F and G to show the top 50 transcription factors. We also appreciate the reviewer’s suggestion to further investigate whether Atf4 directly acts on Ddit3 in the context of APAP toxicity. While this still shows a less pronounced midlobular enrichment for Atf4 compared with Ddit3, we sought additional evidence for a functional Atf4-Ddit3 link. In primary hepatocytes treated with APAP, we observed nuclear co‑localization of Atf4 and Ddit3 (Author response image 3A) and increased nuclear protein levels of both factors (Author response image 3B), supporting their potential cooperative role. We agree that direct analysis of AAV‑Atf4 samples generated for Figure 5 would provide more definitive evidence; unfortunately, co‑staining for Atf4 and Ddit3 on those tissue sections didn’t work well.

      Author response image 3.

      Subcellular localization of Atf4 and Chop in primary hepatocytes following APAP treatment. (A) Immunofluorescence staining of Atf4 and Chop in primary hepatocytes treated with 10 mM APAP for 6 hours or left untreated (UT). Nuclei were counterstained with DAPI. Scale bar as indicated. (B) Primary hepatocytes were treated with 0, 5, or 10 mM APAP for 6 hours. Cytoplasmic and nuclear fractions were isolated and analyzed by western blot. Lamin B1 and α-Tubulin were used as markers for the nucleus and cytoplasm, respectively

      (6) Related to the previous question, the ATF4 immunostaining in F5A doesn't look convincing, with many brown pigments appearing to be outside of the nucleus.

      We thank the reviewer for this helpful comment. To better demonstrate ATF4 nuclear localization, we have added enlarged insets of the original representative images in revised Figure 5A. These magnified views more clearly show nuclear ATF4 staining after APAP treatment, addressing the concern about extranuclear signal.

      (7) It is not ruled out that AAV expression of Atf4 or Btg2 reduces hepatocyte sensitivity to APAP by affecting the expression of the Cyps needed for activation. In other words, does AAV-Atf4 or AAV-Btg2 change the expression of any of the Cyps relevant to APAP in the 3 weeks before APAP application (F5B)?

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      (8) It is laudable that the authors tried to extend their findings to humans by using snRNA-seq data from a published study (line 391), but it is unclear why they didn't analyze all 10 patients in that study but instead focused on 2 and stated that this small sample number prevented drawing definitive conclusions and could therefore only be mentioned in the discussion.

      We thank the reviewer for this clarification. The analysis mentioned in line 391 originally referred to spatial transcriptomics (ST) data from two ALF patients, not snRNA-seq. For the snRNA-seq dataset, we analyzed all 10 patients, but snRNA-seq lacks spatial resolution and cannot reliably assign zonal identity. We stipulate that snRNA-seq requires viable cells and thus likely excludes necrotic/peri-necrotic areas. Therefore, direct zonal comparison with our ST data was not possible. We have now clarified this in the revised manuscript (Revised manuscript, page 18, line 510-519).

      Reviewer #3 (Public Review):

      The main concern is that the overexpression of ATF4 and DDIT3 is causing reduced cell death and damage by APAP. This makes it harder to understand if these genes are truly increasing survival or if they are just reducing the injury caused by APAP. It may be better to perform overexpression immediately after, or at the same time as APAP delivery. Alternatively, loss-of-function experiments using AAV-shRNAs against these targets could be useful.

      We thank the reviewer for raising this important point. We agree that overexpression prior to APAP administration leaves open the question of whether the observed protection reflects true cytoprotection or simply reduced initiation of injury. To address this, we pursued loss‑of‑function approaches. Due to their very low basal expression, AAV‑shRNA‑mediated knockdown of endogenous Atf4 and Ddit3 proved inefficient. We therefore targeted Btg2, a downstream mediator of Ddit3 that inhibits proliferation. Knockdown of Btg2 resulted in a significant increase in APAP‑induced liver injury, as evidenced by elevated ALT levels and expanded necrotic areas (Revised Figure 6K-N). These results indicate that the ATF4‑DDIT3‑BTG2 axis limits hepatocellular damage, consistent with a protective role. We have clarified this point in the revised manuscript (page 15, line 407-414)

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Clarify how zones were defined when necrosis disrupted pericentral areas. Provide marker validation across time and whether necrotic spots are excluded or not from zonal analysis.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g., Alb, Mup20, Cyp2f2, Pck1, Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2, Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g., Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) Test whether the ISR-Btg2 program applies in other models; even targeted validation via qPCR and IF would be valuable.

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Proliferation quantification in liver sections in Figure 1: how to define the zones and why, at the basal level, there is a high proliferation rate in the mid zone? From Figure 1B-C, all three zones showed decreased hepatocyte proliferation, although the mid zone had a higher baseline. Will the mid-zone stand out by converting to the fold change of Ki-67+ hepatocytes decrease?

      We thank the reviewer for these insightful comments. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x<sup>2</sup> + z<sup>2</sup> – y<sup>2</sup>) / (2z<sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885). Consistent with previous reports (PMID: 33632817; PMID: 33632818), we observed a higher baseline proliferation rate in the mid zone, where approximately 70% of proliferating hepatocytes reside under basal conditions, compared to 10% in zone 1 and 20% in zone 3. However, when analyzing the fold change in Ki-67+ hepatocytes, only Mid and PC region showed significant difference in Ki-67+ hepatocytes, other zones showed no significant differences (as shown in the fold-change results in Author response image 1), indicating that the mid zone does not stand out in the fold change analysis. See Author response image 1.

      (4) The authors need to strengthen the causal chain with rescue experiments, e.g., Atf4/Chop overexpression and Btg2 knockdown. Link proliferation suppression to survival/ALT directly.

      We thank the reviewer for these constructive comments. Besides existing data from Figure 5 (Atf4 overexpression), we included Btg2 knockdown data in the revised Figure. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup> hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N) (Revised manuscript, page 14, lines 399–406).

      (5) Transduction efficiency, distribution, and expression levels via the AAV overexpression need to be quantified. Key CYP genes in the APAP metabolic pathway need to be assessed to exclude confounds.

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (6) The authors claim that the requirement of the Atf4/Chop at the early stage of APAP injury protects hepatocytes from proliferation for survival. What is the consequence if we remove the protective mechanism?

      We thank the reviewer for this insightful question. In our model, early induction of Atf4 and Chop functions as a cell survival checkpoint. Removal of this protective mechanism is predicted to result in two deleterious outcomes: (1) Acute exacerbation of necrosis due to the inability of hepatocytes to manage stress-induced bioenergetic demands, and (2) Impaired long-term regeneration due to depletion of the surviving cell pool. We directly tested the acute prediction (< 24 h) in Author response image 4. We deleted Ddit3 specifically in hepatocytes. Initial attempts using AAV-CasRx failed due to negligible baseline Atf4/Chop expression in healthy liver, preventing effective knockdown. We therefore generated hepatocyte-specific Ddit3 knockout mice (Alb<sup>∆Ddit3</sup>; Author response image 4B). Immunohistochemistry confirmed APAP-induced Chop induction occurs primarily in the centrilobular zone by 6 h (Author response image 4A). Following a two-dose APAP regimen (Author response image 4C), Alb<sup>∆Ddit3</sup> mice displayed significantly larger areas of centrilobular necrosis compared to Ddit3<sup>fl/fl</sup> controls (Author response image 4D; **p < 0.01). Thus, hepatocyte-intrinsic Chop limits acute APAP injury, consistent with its proposed early protective role.

      Author response image 4.

      Hepatocyte-specific deletion of Ddit3 exacerbates APAP-induced liver injury. (A) Immunohistochemical staining of Chop in liver sections at 0,3 and 6 h post-APAP. Red arrows indicate Chop-positive hepatocytes. Scale bar = 50μm. Quantification of zonal distribution of Chop-positive cells in liver sections at 6 h post-APAP is conducted . The statistic is the percentage of Chop-positive hepatocytes in each layer over the total number of Chop-positive hepatocytes. n=3 mice. (B)The construction, genotyping strategy and genotyping results of Alb<sup>∆Ddit3</sup> mice. P: positive control; WT: Wild-type; Neg: Blank control(ddH<sub>2</sub>O). (C) Schematic figure illustrating the experimental strategy for the administration of two doses of APAP to Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice. (D) H&E staining showing liver morphology from Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice at 6 h post-second dose of APAP. Injured area is outlined by black dashed lines. Scale bars = 200 μm. The percentage of injury area is quantified. n = 3- 4 mice/group. Data are represented as means ± SD; *p < 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001; ns, not significant.

      (7) Is there any human relevance to the sensitivity of APAP injury regarding the Atf4/Chop axis?

      We thank the reviewer for this insightful comment. During our study, we analyzed a spatial transcriptomics dataset from APAP patients. In one of two analyzed patients, mid-zone hepatocytes exhibited transcriptional signatures remarkably consistent with our murine findings, including: (1) upregulation of Atf4-Chop pathways, and (2) downregulation of cell proliferation genes (Author response image 5). This suggests that this axis may also be involved in the response to APAP injury in humans. However, given the limited sample size, definitive conclusions cannot be drawn at this stage. We have now included this point in the Discussion section (Revised manuscript, page 18, line 510-519).

      Author response image 5.

      Spatial transcriptomics (GSE223561) reveals zonal gene expression changes in APAP patients. Heatmap of ISR, cell death, and cell cycle gene expression across zonal regions in healthy versus APAP‑treated human livers. 

      (8) Several IHC stainings have a weak signal and need inserts to zoom in for a clear view of the positive signals. Figure 5A, E, G, and Figure 6D, F.

      We thank the reviewer for this observation. We agree that the immunostaining signals for several target genes are relatively weak, which reflects their low endogenous expression levels. To address this, we have included higher-magnification insets in the indicated panels (Revised Figure 5A, E, G and Figure 6D, F) to show the positive signals.

      Reviewer #2 (Recommendations for the authors):

      (1) What is the functional classification of DEG in F2A based on? GO terms?

      We thank the reviewer for this constructive question. The functional classification of differentially expressed genes (DEGs) in F2A is based on Gene Ontology (GO) terms. For each DEG, we retrieved its associated GO annotations across the three main categories (biological process, cellular component, molecular function). In cases where a gene was assigned multiple GO terms, we prioritized the most representative or significantly enriched term for functional interpretation. This clarification has been incorporated into the revised figure legend and the according GO number has been included in the figure.

      (3) The rationale for focusing on CHOP is not clear because Ddit3 is not shown in the spatial transcriptomics in F2A and is not significant in F2B, contradicting what is stated in line 206.

      We thank the reviewer for raising this important point. We apologize that Ddit3 was missing from the original figure. In the revised manuscript, we have included an updated version of Figure 2A, which now shows that Ddit3 is indeed one of the differentially expressed genes (DEGs) in the Mid zone at both 3 and 6 hours post-APAP. We agree with the reviewer that, as shown in Figure S1G (previous Figure 2B), Ddit3 did not reach statistical significance, due to its relatively low expression level in that analysis. Nevertheless, when we examined transcription factor (TF) activity in the Mid zone during early AILI, Ddit3 and Atf3 ranked as the top two most highly expressed TFs among the top ten with the highest activity, whereas Atf4 ranked seventh (Revised Figure 4B and Figure S3B). Given that Ddit3 frequently co-worked with Atf4 and that the Atf4–Ddit3 axis plays a well-established role in cellular stress adaptation, we considered this pathway to be biologically relevant and worthy of further investigation.

      (3) The term "redistribution" used in line 197 to describe the expression of Cyp2e1 and other Cyps in the midlobular zone seems inappropriate, considering that they just continue to be expressed there, whereas pericentral hepatocytes are dying in F3B; the same applies to "Gene Expression Shift" in F3H.

      We thank the reviewer for this important clarification. We have revised the text (Revised manuscript, page 9, line 234-236) to state that selective loss of Cyp‑expressing pericentral hepatocytes leads to the mid‑zone becoming the primary site of residual Cyp activity. The figure label has been changed from “Gene Expression Shift” to “Peri‑necrotic Cyp retention” and the legend now explicitly notes that this is an apparent zonal shift due to necrosis, not active redistribution.

      Reviewer #3 (Recommendations for the authors):

      (1) Please do not use abbreviations like AILI. This makes the paper more difficult to read.

      We thank the reviewer for pointing this out. We have replaced AILI with the full term “APAP-induced liver injury” to ensure easiness for readers.

      (2) It will be important to clarify how pericentral, mid, and periportal were defined. In Figure 1, it appears that some of the pericentral hepatocytes that are Ki67 positive are quite mid-zonal. It would be important to have rigorous definitions for the location determination.

      We thank the reviewer for this constructive comment. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x <sup>2</sup> + z <sup>2</sup> – y <sup>2</sup>) / (2z <sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885).

      We thank the reviewers for their rigorous critique again. We thank eLife for fostering an environment of fairness and transparency that enables authors to communicate openly and present their data honestly.

    1. eLife Assessment

      This study investigates the role of Interleukin-2-inducible T cell kinase (ITK) deficiency in autoimmune lung injury using a pristane-induced pulmonary hemorrhage (PH) model, suggesting that ITK-deficient regulatory T cells (Tregs) restrict severe tissue pathology. The work represents a valuable addition to the fields of autoimmunity, inflammation, and T-cell biology in the lung. However, the experimental evidence supporting the underlying cellular and molecular mechanisms and the integration of foundational background literature to provide the necessary context are incomplete.

    2. Reviewer #1 (Public review):

      In this study, Hossain et al. investigated the role of Interleukin-2-inducible T cell kinase (ITK) in autoimmune lung injury, demonstrating that ITK-deficient (Itk-/-) mice are protected against pristane-induced pulmonary hemorrhage (PH). The authors suggest that this protection correlates with a significant remodeling of the T cell compartment in Itk-/- mice, including increased frequency of memory-like CD4+ and CD8+ T cells (CD44⁺CD62L⁺) as well as higher frequency of Treg populations. Furthermore, adoptive transfer of ITK-deficient Treg isolated from injured ITK-deficient mice confers protection against pulmonary hemorrhage in WT recipients.

      Strengths:

      The adoptive transfer of wild-type and Itk-/- Treg populations demonstrates that ITK-deficient Treg can actively rescue pre-existing lung injury and reverse systemic secondary metrics like proteinuria in wild-type recipients, providing proof-of-concept validation for the therapeutic utility of the ITK-Treg axis.

      Weaknesses:

      A primary limitation of this manuscript is its omission of foundational literature from the Schwartzberg and Littman laboratories, which originally established the indispensable role of IL-2-inducible T-cell kinase (ITK) in proximal T-cell receptor (TCR) signaling dynamics and thymic lineage commitment. Because classic studies demonstrate that ITK is a critical regulator of thymic T cell development and cellular proliferation (PMID: 8777721, 10213685), the authors' claim that "these findings indicate that ITK deficiency skews the T cell compartment toward a memory-like state, establishing a distinct immune baseline that may favor protective and regulatory responses over pathogenic inflammation" is not substantiated by evidence and requires more robust validation.

      The exclusive reliance on splenic immunophenotyping is a major limitation, as it fails to capture the local cellular dynamics within the primary organs of injury (the lung and kidney). Evaluating canonical and non-canonical Treg expansion solely in the spleen overlooks the distinct functional programming of tissue-resident subsets. The authors should extend their characterization of regulatory T cell compartments directly to the lungs and draining lymphoid structures.

      More importantly, the authors overlook key historical publications that explicitly established ITK as a negative "rheostat" or gatekeeper for regulatory T cell (Treg) differentiation. Specifically, Huang et al. (PMID: 25063868) previously demonstrated that Treg abundance is inversely correlated with ITK expression, and that ITK activity serves as a vital negative tuner of IL-2-driven Foxp3⁺ Treg expansion. Since it is already well-established that suppressing or deleting ITK promotes Treg accumulation and function, and that these cells are intrinsically vital to suppressing systemic autoimmunity, it is unclear how these findings expand upon our existing mechanistic understanding of ITK regulatory biology.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Hossaim and colleagues investigate the role of the ITK kinase in modulating inflammation in a pristane-induced lung hemorrhage model. Using a germline ITK KO mouse, they report that loss of ITK skews the T cell compartment toward a memory-like state, expanding Tregs, and conferring protection against alveolar hemorrhage, inflammatory monocyte recruitment, proteinuria, and systemic cytokine elevation. They further show that transfer of ITK-deficient Tregs into wild-type hosts with established disease attenuates injury and shifts the cytokine balance toward resolution, and that ITK-deficient Tregs carry a transcriptional signature enriched for OXPHOS, mTORC1, MYC, and cell-cycle programs. While these observations are interesting for the development of potential immunotherapies, there are several issues with the methodological approach that support the authors' claims, tempering my enthusiasm for this manuscript.

      Strengths:

      (1) The clinical motivation and potential targeted therapies are relevant.

      (2) The murine phenotype seems robust.

      Weaknesses:

      (1) All loss-of-function experiments are from a global ITK knockout. This is a major limitation and weakness of this study. The protection observed in the intact knockout, therefore, cannot be attributed to Tregs specifically. The Treg-intrinsic claim rests almost entirely on a single adoptive-transfer experiment. In order to show that this effect is Treg-specific, the authors would need to generate a Treg-specific ITK-deficient mouse

      (2) In their sufficiency experiment (adoptive Treg cell transfer), donor and/or host cells are not congenically marked, so persistence, lung trafficking, and in vivo expansion of transferred Tregs are not demonstrated.

      (3) The authors claim that ITK-deficient Tregs possess enhanced metabolic fitness. This conclusion is based on transcriptional profiling of isolated splenic Tregs from unchallenged mice, yet it concerns lung protection during active disease. A disease-state and ideally lung-relevant transcriptome would more directly support the mechanistic narrative. Additional functional validation would be needed (Seahorse assay, mitochondrial mass/potential, etc). Some of these GSEA programs enriched in ITK-deficient Tregs could reflect a more general proliferative signature.

    4. Reviewer #3 (Public review):

      Summary:

      Hossain et al. investigate the role of ITK as a central regulator of autoimmune lung injury. They used ITK-deficient mice and the pristane-induced pulmonary hemorrhage (PH) model to show that ITK deficiency confers protection against PH. The adoptive cell transfer experiment suggests a possible role for altered Treg cells in ITK-deficient mice in regulating the inflammatory response in the lungs of pristane-injected mice. This study shows that targeting the ITK axis may be beneficial by reducing systemic inflammatory injury that contributes to poor outcomes in PH.

      Strengths:

      This study highlights the importance of ITK in regulating pulmonary hemorrhage. The enrichment of Treg cells is known to confer protection in autoimmunity-mediated alveolar damage. However, ITK's involvement in regulating Treg cell function is interesting and could be explored as a novel therapeutic approach for chronic inflammation.

      Weaknesses:

      The novelty of this study lies in the association between ITK-deficient Tregs and pulmonary hemorrhage in autoimmunity. The weakness of the manuscript is the lack of sufficient experiments to support the claim that ITK-deficient mice show protection specifically mediated by Treg cells, and to demonstrate that ITK-deficient Treg cells are more efficient than WT Treg cells in regulating other immune cells that drive pulmonary damage. The authors performed all the experiments in ITK global knockout mice, in which not only T cells but all other cell types are deficient in ITK. Furthermore, they have not performed any functional analysis to demonstrate the functional differences between WT Treg and ITK-deficient Treg cells, undermining the novelty of this study.

    1. eLife assessment

      The authors present a valuable open-source tool for three-dimensional analysis of dissected slices of human brains including 3D reconstruction and high-resolution 3D segmentation. Convincing evidence is provided based on experiments on both real and synthetic data. This tool would be of use to researchers in the neuropathology and neuroimaging field.

    1. eLife assessment

      The study presents a tool for searching molecular dynamics simulation data, making such data sets accessible for open science. The authors provide convincing evidence that it is possible to identify noteworthy molecular dynamics simulation data sets and their analysis can produce valuable information.

    1. eLife Assessment

      This valuable manuscript by Alonso-Caraballo et al is a novel piece of work that examines the impact of oxycodone self-administration on neural plasticity within paraventricular thalamic (PVT) to nucleus accumbens shell (Shell) pathway - two regions shown to play a key role in cue-induced drug seeking on their own - and whether this plasticity varies based on abstinence period and biological sex. Data show that a clinically relevant long-access model of self-administration promotes dependence in both male and female rats and provide compelling data that when compared to current literature indicate that craving-induced relapse for opioids may develop faster and may be more pronounced in females compared to males. In addition to these behavioral findings, the authors provide the first evidence that glutamate signaling within the PVT-to-Shell pathway is selectively strengthened at the output medium spiny neurons by opioids following protracted, but not acute abstinence. These data highlight a potential role for these adaptations in relapse behavior and identify a potential therapeutic target during abstinence to reduce relapse risk in abstaining individuals.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have made minor revisions to address the comments raised in the previous round of review.]

      Summary:

      This manuscript by Alonso-Caraballo et al, is a novel piece of work that examines the impact of oxycodone self-administration on neural plasticity within paraventricular thalamic (PVT) to nucleus accumbens shell (Shell) pathway - two regions shown to play a key role in cue-induced drug seeking on their own, and whether this plasticity varies based on abstinence period and biological sex.

      Strengths:

      The authors show using a clinically relevant long-access model of opioid self-administration promotes dependence and acute withdrawal in both male and female rats. During subsequent cue-induced relapse tests at 1 or 14-days following the conclusion of self-administration, data show that while both male and females demonstrate drug-seeking behavior at both time points, females show a further elevation in responding on day 14 versus day 1 that is not observed in the males. When accounting for past work showing elevations in drug seeking in males after 30 days, these data indicate that craving-induced relapse for opioids may develop faster and may be more pronounced in females compared to males.

      These behavioral findings were paralleled by use of ex vivo acute slice electrophysiology and circuit-specific ex vivo optogenetics to examine the impact of oxycodone self-administration on synaptic strength within the paraventricular thalamus (PVT) to nucleus accumbens shell (NAcSh) pathway(s). Data support a time-dependent but sex independent strengthening of glutamatergic signaling at PVT-to-NAcSh medium spiny neurons (MSNs) that is only present following a relapse test at 14 days post abstinence in males versus females, providing the first evidence that opioid self-administration and/or cue-induced drug-seeking augments this pathway. Using an extensive set of physiological measures, the authors show that this increased synaptic strength reflects a upregulation of presynaptic release probability. Further, this upregulation of excitatory signaling aligned temporally with an increase in MSN excitability, as assessed by increases in action potential firing frequency. Finally, the authors provide the first evidence that similar to other inputs to the NAcSh, PVT projections innervate both MSN as well as local interneurons, promoting a GABA-A specific feedforward inhibitory circuit. Interestingly, unlike direct excitatory inputs to MSNs, no changes were observed ostensibly within this feedforward circuit, highlighting a selective enhancement of excitatory drive and output of MSNs with protracted abstinence.

      Overall, these data highlight a potential role for heightened synaptic strength within the PVT-NAcSh pathway in cue-induced relapse behavior during protracted abstinence and identify a potential therapeutic target during abstinence to reduce relapse risk in abstaining individuals.

      Weaknesses:

      Overall, the experimental approach and data provided appear rigorous and support their overall conclusions and achieve their goal of understanding how opioid self-administration impacts synaptic strength within the PVT-NAcSh pathway. Although not undermining these data, there are a few potential weaknesses that reduce the impact of the work. For example, the inability to directly assess whether cue-induced drug-seeking is in fact augmented compared to daily intake during self-administration in the maintenance face only permits the authors to denote that reexposure to cues and the context is sufficient to promote active lever pressing without demonstrating whether seeking behavior is in fact elevated further during a cue test. This is notably understandable as drug available sessions were 6-hours versus a 1hour relapse test. Importantly, it is clearly demonstrated that drug seeking is higher on average in female mice after 14 days versus 1 day.

      With regard to interpretation of electrophysiology findings, the lack of inclusion of an abstinence only group does not permit interpretations to parse out whether observed increases in synaptic strength (or the lack of) reflect abstinence or an interaction between abstinence period and re-exposure to the operant chamber, as slices were taken 30-45 min post relapse test. While much literature has shown that drug induced adaptations in the NAc requires a post drug period for plasticity to measurably emerge, studies have also shown that re-exposure to heroin-associated cues following abstinence seemingly "reverses" increases in cell excitability in prelimbic-NAc pyramidal neurons (Kokane et al., 2023) and that depotentiation of morphine-induced increases in synaptic strength in the NAc shell can be depotentiated by drug re-exopsure -- an effect also observed with cocaine re-exposure (Madayag et al., 2019). Notably, the lack of effect at 14 but not 1 day supports the likelihood that the relapse test does not in fact influence the plasticity within the PVT-NAcSh circuit.

      While the lack of effect on AMPAR:NMDAR ratio and rectification indices do support the notion that enhanced EPSC amplitudes in input-output curves do not reflect a change in AMPAR subunit expression (i.e., increased GluA2-lacking receptors that exhibit inward rectification at depolarized potential) nor a change in postsynaptic sensitivity to glutamate, without direct assessment of AMPAR-specific and NMDAR-specific input-output curves, it doesn't definitively exclude the possibility that both AMPA and NMDA receptor currents are being upregulated, thus negating an observable change in postsynaptic strength.

      Overall, these findings provide novel insight into how the PVT-NAcSh pathway is altered by opioid self-administration and whether this is unique based on abstinence period and sex. Importantly, these were the primary objectives stated by the author. Data highlight a potential role for the observed adaptations in relapse behavior and identify a potential therapeutic target during abstinence to reduce relapse risk in abstaining individuals. However, it should be noted that no causal link is demonstrated without experiments to reduce/prevent relapse.

      Comments on previous revisions:

      The authors addressed previous concerns brought up, specifically by clarifying data interpretation as well as text modifications related to potential caveats of these interpretations.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting paper from Alonso-Caraballo and colleagues that examines the influence of opioid use, acute and prolonged abstinence, and sex on cue-induced relapse and paraventricular thalamus (PVT) to nucleus accumbens shell (NAcSh) medium spiny neurons circuit physiology. The study presents a valuable finding that following prolonged, but not acute abstinence from oxycodone self-administration, female rodents exhibit higher relapse rates to drug paired cues. Additionally, the study presents the useful finding that prolonged abstinence increased PVT-NAcSh MSN synaptic strength in both sexes, an effect that is likely due to presynaptic adaptations. While the evidence to support these two findings is solid, further experiments are required to determine the functional role of the PVT-NAcSh MSN circuit in relapse following prolonged oxycodone abstinence, and the mechanism underlying the heightened relapse vulnerability in females in this model of opioid use disorder.

      Strengths:

      The paper is interesting, well written and presented, and the experiments are well designed and conducted. The revised analysis of spike count data that models the hierarchical structure of the data is appropriate to overcome low animal numbers and the potential for oversampling. The authors are transparent in reporting the results related to this analysis in figure 5 and acknowledge the study is underpowered to confirm the trend of increased intrinsic excitability in male MSNs following prolonged oxycodone analysis.

      Impact:

      The topic is of interest to the field of substance use disorders and gives solid evidence for the need to consider targeted therapeutics aimed at relapse prevention in opioid use disorder.

    4. Reviewer #3 (Public review):

      Summary:

      Alonso-Caraballo et al. use behavioral testing and ex vivo patch-clamp electrophysiology combined with circuit-specific optogenetic stimulation of PVT terminals to examine how oxycodone self-administration and abstinence duration shape cue-induced relapse and PVT-NAcSh synaptic transmission in male and female rats. In the revision, the authors reanalyzed intrinsic excitability using nested hierarchical GLMMs, acknowledged the low power in the male prolonged-abstinence group, and expanded the discussion of relevant PVT-NAc literature. These changes improve the manuscript. That said, most of the revisions are textual and the main experimental gap remains. Both sexes show increased oxycodone seeking compared to saline at 14 days, but only females show a time-dependent incubation from 1 to 14 days, and the PVT-NAcSh synaptic strengthening is the same in both sexes. Nothing in the revision brings those two observations closer together. The excitability data also come from NAcSh MSNs with no confirmation of PVT connectivity, which limits what circuit-specific conclusions can be drawn. The study is a solid characterization of abstinence-related synaptic changes in this pathway, but some of the conclusions still go further than the data allow.

      Strengths:

      The behavioral characterization is thorough and well-executed, covering self-administration, somatic withdrawal, and cue-induced relapse across two abstinence durations in both sexes. The sex-specific escalation in oxycodone seeking from 1 to 14 days in females but not males is a clear and compelling finding. The use of circuit-specific ex vivo optogenetics to isolate PVT terminal inputs onto NAcSh neurons is a genuine methodological strength, and the demonstration of feedforward inhibitory recruitment through local GABAergic interneurons adds meaningful novelty to the circuit characterization. The reanalysis of intrinsic excitability using nested hierarchical GLMMs appropriately accounts for the non-independence of cells recorded within the same animal and is a real improvement over the original approach. The expanded discussion of prior PVT-NAc work, particularly the more accurate treatment of Keyes et al. (2020) and Paniccia et al. (2024), better situates the findings within the existing literature.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1:

      I recommend that the title be changed to not focus on sex differences to avoid misunderstanding.

      We thank Reviewer #1 for this suggestion and agree that the original title could create a misleading impression. We have updated the title from "Sex-specific behavioral and thalamo-accumbal circuit adaptations after oxycodone abstinence" to "Thalamo-accumbal circuit adaptations following extended oxycodone abstinence" to more accurately reflect the scope of the findings.

      The authors should also address the lack of difference physiologically compared to the behavior as a caveat more clearly in the discussion.

      We thank the reviewer for this important suggestion. We have revised the Discussion to explicitly address this dissociation. Specifically, we added the following to the PVT-NAcSh synaptic strength section: " The absence of sex differences in PVT-NAcSh synaptic measures suggests that this pathway, as characterized here, represents a shared neuro-adaptation to prolonged oxycodone abstinence rather than a substrate for the heightened relapse vulnerability observed in females. The mechanisms driving sex-specific relapse likely involve additional circuit elements, such as sex hormone-dependent modulation, upstream inputs, or cell-type specific plasticity." This point is also summarized in the abstract.

      Reviewer #2:

      A major weakness of this study is the disconnect between the behavioral and neurophysiological data reported. While a striking sex difference in relapse-like behavior is observed, there are no statistically significant sex differences in any of the neurophysiological data reported. Moreover, without an experiment to functionally test the role of the PVT-NAc projection in relapse-like behavior following prolonged oxycodone, these two arms of the study seem divorced.

      We respectfully disagree with the characterization that the behavioral and neurophysiological data are "divorced." The two arms of the study converge on a consistent and meaningful finding: PVT-NAcSh synaptic strength increases specifically after prolonged abstinence, this is the same time point at which enhanced cue-induced relapse is observed in both sexes. The absence of sex differences in synaptic measures does not weaken this convergence; it refines it by suggesting that circuit-level potentiation is a shared neuro-adaptation, while the sex-specific behavioral phenotype likely reflects additional modulatory mechanisms acting on this shared substrate. We have revised the Discussion to explicitly address this dissociation, as noted in our response to Reviewer #1 above. We acknowledge that functional manipulation of the PVT-NAcSh circuit would be required to establish causality, and we state this clearly in the manuscript.

      In the introduction the authors state they aim to test the hypothesis that increased synaptic strength in PVTNAcSh projections are necessary for drug-seeking. This study does not include the required experiments to test this hypothesis.

      We have revised the relevant section in the Introduction to accurately reflect the scope of our study: " We aimed to determine whether synaptic strength in PVT-NAcSh projections is affected following oxycodone abstinence and whether such changes are associated with cue-induced relapse and drug-seeking. Additionally, we examined whether there are sex-specific differences in either cue-induced relapse or PVT-NAcSh synaptic transmission after either 1 (acute) or 14 (prolonged) days of forced abstinence. Our results demonstrate that sex-specific enhancement in cue-induced relapse emerges after prolonged abstinence but not during acute abstinence from oxycodone self-administration. Although both males and females show increased cue-induced relapse after prolonged abstinence, females exhibited a greater relapse rate compared to males. Both sexes showed similar increases in PVT-NAcSh synaptic strength after prolonged abstinence, while synaptic strength was not altered after acute abstinence compared to saline controls. Together, these findings reveal a time-dependent increase in PVT-NAcSh synaptic strength and a sex-specific effect of prolonged abstinence on cue-induced relapse, while synaptic enhancements after prolonged abstinence were not sex-specific." This revision avoids implying a necessary or causal role for the circuit, which we did not test.

      Reviewer #3:

      The PVT-NAcSh synaptic strengthening after prolonged abstinence is statistically indistinguishable between sexes, while females but not males show a time-dependent escalation in oxycodone seeking from 1 to 14 days of abstinence. The Discussion proposes hormonal modulation or differences in upstream inputs as possible explanations, but none of these are tested and the gap is left unresolved.

      We agree that the mechanistic basis of the behavioral sex difference remains an open question that the current study does not resolve. As noted in our response to Reviewer #1, we have revised the Discussion to explicitly acknowledge this dissociation and to clarify that PVT-NAcSh synaptic strengthening represents a shared neuro-adaptation rather than a mechanism specific to the female behavioral phenotype. We maintain that identifying this dissociation is itself a scientifically meaningful finding.

      The intrinsic excitability recordings come from NAcSh MSNs with no confirmation that those neurons receive direct PVT input, which was raised in the original review, acknowledged in the revision, and not experimentally addressed.

      We have added the following clarification to the excitability section of the Discussion: “It should be noted that the intrinsic excitability recordings were designed to characterize general properties of NAcSh MSNs following oxycodone abstinence, independent of their synaptic inputs. As such, the excitability data should be interpreted as reflecting changes in the NAcSh MSN population broadly rather than in PVT-connected neurons specifically. The standing theory suggests that MSN excitability decreases as a homeostatic response to increased glutamatergic input [23,43,58]. Our data do not support a compensatory decrease in excitability in either sex at either abstinence time point”. These recordings were never intended to be circuit-specific; the experiment was designed to characterize NAcSh MSN excitability at the population level, which is a valid and informative question in its own right.

      The male prolonged-abstinence excitability trend has approximately 20% statistical power and is non-significant, yet the Discussion interprets it as a potential neuro-adaptation that could facilitate signal flow through the PVT-NAcSh circuit and contribute to relapse, which goes well beyond what the data support.

      We have revised the relevant Discussion text to ensure the male excitability trend is interpreted appropriately. The revised text now reads: "In males, a non-significant trend toward increased excitability was observed after prolonged abstinence, with a large effect size (Cohen's d = 1.18); however, given that this group was substantially underpowered (approximately 20% power), this finding should be interpreted with caution and cannot be taken as evidence of a neuro-adaptation. Whether this trend, if confirmed in future studies with larger cohorts, reflects a broader MSN population response or is specific to PVT-connected neurons remains an open and interesting question." The speculative mechanistic interpretation previously present in this section has been removed.

      The failure to distinguish between D1 and D2 MSNs remains a significant limitation given that cell-type specific plasticity at PVT-NAc synapses has been shown to be directly relevant to opioid seeking in prior work.

      We agree that distinguishing between D1 and D2 MSNs would provide important mechanistic insight, and we acknowledge this explicitly as a limitation and a future direction in the Discussion. The use of transgenic Cre rat lines for cell-type-specific recording in a self-administration model requires significant additional infrastructure and was beyond the scope of the present study. This is precisely the direction our laboratory is currently pursuing, and the present findings provide empirical motivation for those experiments.

      The Conclusion builds a mechanistic framework around D2 MSNs, PV interneurons, and D1 MSNs that is drawn from studies using different drugs or experimental designs, and none of these cell-type-specific mechanisms are tested in the present experiments.

      We thank the reviewer for this important critique. We have revised the opening of the Conclusion to clarify that the cell-type-specific framework is grounded in prior literature and represents a hypothesis for future investigation rather than a conclusion drawn from the present data. The revised text now reads: " When considered alongside prior work, our findings highlight the need to examine the anatomical and cell-type specific organization of PVT inputs to the NAcSh in the context of opioid relapse. Based on existing literature, PVT projections onto D2 MSNs and PV interneurons may contribute to relapse vulnerability, while adaptations involving D1 MSNs may underlie incubation of craving, though these mechanisms remain to be directly tested in the oxycodone self-administration model used here". We believe this framing accurately represents the relationship between our findings and the broader literature without overstating what the present data demonstrates.

    1. eLife Assessment

      This important study addresses the role of sphingolipid metabolism in maintaining endolysosomal membrane integrity and its impact on tau pathology in Caenorhabditis elegans and human cell culture models. The findings are convincing, and the proposed mechanisms are conceivable. The experimental evidence supports the conclusions of the study. The work will be of broad interest to cell biologists and biologists working on Alzheimer's disease and related proteinopathies.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, Tittelmeier et al. explored the role of sphingolipid metabolism in maintaining endolysosomal membrane integrity and its downstream effects on tau aggregation and toxicity, using both worms and human cell models. The authors showed that knockdown of sphingolipid metabolism genes reduced endolysosomal membrane fluidity, as revealed by FRAP and C-Laurdan imaging, leading to increased vesicle rupture. Furthermore, tau aggregates accumulated in endolysosomes and exacerbated membrane rigidity and damage, promoting seeded tau aggregation, likely by enabling tau seed escape into the cytosol. Importantly, unsaturated fatty acid supplementation restored membrane fluidity, suppressed tau propagation, and alleviated neurotoxicity in C. elegans. These findings provide insight into how lipid dysregulation contributes to tau pathology and highlight membrane fluidity restoration as a potential therapeutic avenue for Alzheimer's disease.

      Strengths:

      The study addresses the connection between sphingolipid metabolism, endolysosomal membrane integrity, and tau pathology, which is a relevant topic in the context of Alzheimer's disease and related tauopathies.

      The use of both C. elegans and human cell models provides cross-species perspectives that help frame the findings in a broader biological context.

      The combination of FRAP and C-Laurdan dye imaging offers a biophysical approach to investigate changes in membrane properties, which is a technically interesting aspect of the study.

      The observation that unsaturated fatty acid supplementation can modulate membrane fluidity and influence tau-related phenotypes adds an element of potential therapeutic interest.

      The study presents multiple experimental approaches to address the proposed mechanism, and efforts were made to examine both membrane behavior and tau aggregation dynamics.

      Comments on revised version:

      I thank the authors for their thorough revisions and detailed responses. All of my previous concerns have been satisfactorily addressed, and I have no further comments.

    3. Reviewer #2 (Public review):

      Tittelmeier et al. investigated the role of sphingolipid (SL) metabolism in the maintenance of endolysosomal vesicle integrity. They find that both impaired SL biosynthesis and degradation in C. elegans decreases the fluidity of endolysosomal membranes and promotes their rupture, while it has little effect on plasma membrane fluidity. Endolysosomal membrane fluidity is also negatively affected in human cells upon knockdown (KD) of a gene (SPHK2) involved in the SL degradation pathway. Aggregated forms of tau in both models (C. elegans and human cells) can also cause rigidification of the endolysosomal membrane, with SL homeostasis disruption having an additive effect, exacerbating endolysosomal rupture. Notably, KD of SPHK2 also increased the formation of tau foci, suggesting that compromised endolysosomal integrity may promote tau aggregation. These data provide a clearer understanding of how genetic manipulation of SL metabolism affects endolysosomal membranes and their rigidification in the context of tau aggregation. Supplementation of polyunsaturated fatty acids (PUFAs), which has a beneficial effect on Alzheimer's patients, improved membrane fluidity and reduced tau propagation in human cells and tau-associated neurotoxicity in C. elegans, suggesting a possible mechanism of action.

      Comments on revised version:

      The authors have:<br /> Corrected editorial errors (Points 1 and 2).

      Clarified the experimental rationale, added new data to rule out alternative explanation, and improved the presentation of the C. elegans model (Point 3).

      Provided experimental evidence and appropriate discussion regarding the specificity and broader physiological context of SL gene knockdown effects (Point 4).

      Overall, the authors' responses are thorough, supported by new data where appropriate, and demonstrate a clear understanding of the concerns raised. All points raised have been satisfactorily resolved.

    4. Reviewer #3 (Public review):

      Summary:

      The authors set off with an analysis of the lysosomal integrity upon knockdown of genes of the sphingolipid metabolic pathway that they identified in a previous work of an RNA screen using a new C.elegans Tau model. They then used cell culture and C.elegans experiments to study the link between lysosomal rupture and Tau propagation.

      Strengths:

      The authors use two complementary model systems and used probes to assess membrane rigidity that allow a quick assessment of the membrane dynamics and offer the opportunity to treat the cells with lipids, RNAi. Tau seeds etc.

      Comments on revised version:

      The authors have addressed the majority of my critical comments and thus I support the manuscript.

      They have still not analysed the knockdown efficiencies of their RNAi experiments. But this is their choice.

      The other publication establishing their Tau model is meanwhile published and there is no disconnect anymore between the model their analysis builds on.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Tittelmeier et al. explored the role of sphingolipid metabolism in maintaining endolysosomal membrane integrity and its downstream effects on tau aggregation and toxicity, using both worms and human cell models. The authors showed that knockdown of sphingolipid metabolism genes reduced endolysosomal membrane fluidity, as revealed by FRAP and C-Laurdan imaging, leading to increased vesicle rupture. Furthermore, tau aggregates accumulated in endolysosomes and exacerbated membrane rigidity and damage, promoting seeded tau aggregation, likely by enabling tau seed escape into the cytosol. Importantly, unsaturated fatty acid supplementation restored membrane fluidity, suppressed tau propagation, and alleviated neurotoxicity in C. elegans. These findings provide insight into how lipid dysregulation contributes to tau pathology and highlight membrane fluidity restoration as a potential therapeutic avenue for Alzheimer's disease.

      Strengths:

      The study addresses the connection between sphingolipid metabolism, endolysosomal membrane integrity, and tau pathology, which is a relevant topic in the context of Alzheimer's disease and related tauopathies.

      The use of both C. elegans and human cell models provides cross-species perspectives that help frame the findings in a broader biological context.

      The combination of FRAP and C-Laurdan dye imaging offers a biophysical approach to investigate changes in membrane properties, which is a technically interesting aspect of the study.

      The observation that unsaturated fatty acid supplementation can modulate membrane fluidity and influence tau-related phenotypes adds an element of potential therapeutic interest.

      The study presents multiple experimental approaches to address the proposed mechanism, and efforts were made to examine both membrane behavior and tau aggregation dynamics.

      We thank the reviewer for this positive assessment of the study.

      Weaknesses:

      In Figure 3, the authors used C-Laurdan imaging to assess membrane fluidity and showed that knockdown of SPHK2, the human ortholog of sphk-1, led to increased membrane rigidity. However, the authors did not co-stain with a lysosomal marker, making it unclear whether the observed effect is specific to lysosomal membranes or reflects general membrane changes. Co-staining with LysoTracker or applying segmentation masks to isolate lysosomal signals would significantly improve interpretation.

      We agree with the reviewer that it is important to isolate lysosomal signals for interpreting the C-Laurdan data. We therefore repeated and extended the C-Laurdan experiments in combination with LysoTracker staining and selectively analyzed LysoTracker-positive regions. These analyses showed pronounced increases in GP values in LysoTracker-positive vesicles after SPHK2 knockdown, supporting the conclusion that SPHK2 depletion increases endolysosomal membrane rigidity. We also performed LysoTracker-based analysis in the tau-fibril and fatty-acid experiments to better assess lysosome-associated membrane properties (see new and updated Figures 3B-E, Figures 5A and B, Figures S3A, B, F, G, Figures S4A-F, and Figures S5A-D for details). The respective Results sections have been revised accordingly.

      Line 173 states that Lipofectamine 2000 increases membrane fluidity based on GP index changes, but this is incorrect. A higher GP index indicates increased membrane order (i.e., reduced fluidity), so the statement should be revised. Additionally, Lipofectamine 2000 can itself alter membrane rigidity, posing a risk of false-positive interpretations. To confirm the role of SPHK2 in this phenotype, the authors should use a CRISPR/Cas9 knockout model instead of relying solely on siRNA transfection, which may be confounded by the delivery reagent. Without lysosomal co-staining and SPHK2 KO validation, the authors cannot conclusively claim that SPHK2 loss affects endolysosomal membrane integrity.

      We thank the reviewer for pointing out the incorrect wording regarding Lipofectamine. A higher GP index indicates increased membrane order/rigidity, not increased fluidity. Since the revised main figure now includes the SH-SY5Y data (Figure 3A-D), in which Lipofectamine alone did not significantly alter GP values (see Figure S3A, B), we removed the misleading statement from the Results.

      We also agree that Lipofectamine can affect membrane properties and therefore needs to be carefully controlled. In all siRNA-mediated experiments, SPHK2 siRNA was compared to a matched control siRNA condition exposed to the same transfection reagent. We state in the Methods that cells were transfected with either SPHK2 or scrambled control siRNA and that the medium was exchanged after 6 h “to minimize lipofectamine impact on the endolysosomal system”. Thus, the effect attributed to SPHK2 KD is assessed relative to the appropriate Lipofectamine-containing control condition.

      Importantly, we have now added lysosomal co-staining to address the reviewer’s concern about compartment specificity. This analysis showed that SPHK2 KD resulted in a pronounced increase in GP values within LysoTracker-positive compartments, demonstrating increased membrane rigidity at lysosomes. Thus, the revised data support the conclusion that SPHK2 KD increases lysosome-associated membrane rigidity, rather than only causing nonspecific effects on other cellular membranes.

      We also clarified the relationship between the current siRNA-based assay and our previous CRISPR inhibition-based analysis. In the revised Results, we now write: “While SPHK2 KD alone significantly increased galectin puncta above the matched control, its effect was more modest than in our previous CRISPR inhibition-based analysis [20]. This difference likely stems from the earlier readout required for the combined siRNA/tau fibril assay, when transient Lipofectamine-associated effects still increased the control background.”

      This addresses why the SPHK2 KD effect appears smaller in the current siRNA/tau-fibril assay than in our previous CRISPR inhibition-based analysis. The previous study, which is now peer-reviewed and published in the journal Autophagy, used a CRISPR inhibition-based strategy to reduce SPHK2 levels, which resulted in a highly significant increase in sfGFP-LGALS3 foci formation compared to the control [1]. Thus, the SPHK2 phenotype is not supported solely by the current siRNA experiment.

      In addition, we sought to genetically validate the RNAi phenotypes using mutant strains. However, mutant strains were not available for all sphingolipid metabolism hits analyzed in this study. We therefore used the sphk-1 mutant strain available at CGC (CZ24969; sphk-1(ju831)) to validate one of the key SL metabolism hits independently of RNAi. The revised manuscript states: “As genetic validation independent of RNAi, we tested an available sphk-1 mutant strain, which also showed a robust increase in hypodermal sfGFP::LGALS3 foci (Figure S1A).” This result supports the conclusion that genetic perturbation of sphingosine kinase activity compromises endolysosomal integrity in vivo.

      Together, the revised manuscript addresses the reviewer’s concerns by correcting the GP interpretation, controlling the siRNA experiments against matched Lipofectamine-treated controls, adding LysoTracker-based lysosome-associated C-Laurdan analysis, relating the current siRNA data to our previous CRISPR inhibition-based analysis, and providing genetic validation for the available sphk-1 mutant.

      The section titled "Fibrillar tau increases membrane rigidity and exacerbates endolysosomal damage" (lines 177-215) requires substantial revision. The narrative jumps abruptly between worms and cell models, making it hard to follow the logic. The use of the F3ΔK281::mCherry strain is introduced without explanation or context. It is unclear whether this strain is relevant to lysosomal membrane rupture, as no reference or justification is provided. The authors should clarify whether this reporter is intended to detect lysosomal membrane permeabilization (LMP). If so, it would be more appropriate to use established LMP reporters, such as lysosome-targeted fluorescent sensors, galectin-based reporters, or dextran leakage assays. Based on the current data in Figure 3G, it is difficult to draw firm conclusions regarding membrane rupture levels.

      We agree that this section required clarification, and we have substantially revised the Results to improve the logic and separation between model systems.

      First, we now introduce the C. elegans reporter strain earlier in the manuscript, in the first Results section. In the revised text, we explain both the tau construct and the actual lysosomal damage reporter: “In this strain, endolysosomal membrane damage is monitored in the hypodermis by expression of human galectin-3 fused to superfolder-GFP (sfGFP::LGALS3). The animals also express an aggregation-prone tau fragment fused to mCherry (F3ΔK281::mCherry) in touch receptor neurons, which is transmitted to the hypodermis, as described previously [20].” We also clarify the principle of the Galectin reporter: “Under steady-state conditions, sfGFP::LGALS3 remains diffusely distributed throughout the cytosol. Upon endolysosomal damage, luminal β-galactosides become exposed and recruit sfGFP::LGALS3 into visible puncta, providing a sensitive readout of vesicle rupture.” Thus, F3ΔK281::mCherry is not the reporter for lysosomal membrane permeabilization; the membrane-damage readout is sfGFP::LGALS3 puncta formation.

      Second, we reorganized the manuscript to separate the human cell experiments from the C. elegans experiments more clearly. The revised section “Fibrillar tau and SPHK2 KD act in concert to exacerbate endolysosomal damage and seeded tau aggregation” now focuses on human cell data. The C. elegans experiments are now presented in a separate section, “Tau transmission sensitizes endolysosomal membranes to sphingolipid perturbations in vivo.” We believe that this revised structure now clearly distinguishes the role of the Galectin reporter from the tau transmission model, separates the human cell and C. elegans data, and avoids the abrupt transitions between model systems noted by the reviewer.

      To support the conclusion that sphingolipid metabolism gene knockdown alters membrane properties, the study would benefit from direct lipidomic analysis. Measuring changes in sphingolipid profiles in both C. elegans and cell models would provide biochemical evidence for the proposed disruption of lipid homeostasis. Given the availability of lipidomics platforms, this type of analysis should be feasible in both worms and human cells and would significantly strengthen the mechanistic claims regarding membrane fluidity and integrity.

      Because we did not perform lipidomics in the present study, we have revised the wording throughout the manuscript to avoid implying that we directly measured lipid composition. Instead, we now refer to “genetic perturbation/disruption of sphingolipid metabolism” or “knockdown of enzymes involved in sphingolipid metabolism” when describing our experimental interventions.

      We agree that lipidomic analyses will be important in future studies to define how perturbation of sphingolipid metabolism changes lipid composition in C. elegans and human cells. However, lipidomics itself would not directly establish which lipid changes causally drive the membrane rigidification observed in our study. Membrane fluidity is a biophysical property determined by the combined composition of the membrane, including lipid abundance, saturation, acyl-chain length, head groups, sterol content, and membrane-associated proteins. Thus, even if lipidomics identified changes in sphingolipid profiles, these changes could not be directly translated into a predictable effect on membrane fluidity without additional biophysical validation, using Laurdan dye imaging or FRAP. Moreover, whole-cell or whole-animal lipidomics would not resolve whether the relevant lipid changes occur specifically at endolysosomal membranes, which are the focus of our study.

      We have now clarified this point in the Discussion. Specifically, we state that “even detailed lipidomics would not by itself identify which lipid changes are responsible for the observed membrane rigidification” and that future lysosome-enriched or organelle-specific lipidomic approaches should be combined with direct manipulation of candidate lipid species, followed by measurements of membrane fluidity and rupture, to determine which lipid changes causally contribute to endolysosomal membrane rigidification. In the present study, we therefore focused on direct quantitative biophysical readouts of membrane properties in C. elegans. We used FRAP of the lysosomal membrane protein LAAT-1::mCherry to assess lateral mobility within lysosomal membranes and showed that knockdown of sphingolipid-metabolism genes increased the time to half-maximal recovery, indicating reduced lysosomal membrane fluidity. Notably, knockdown of genes involved in both sphingolipid biosynthesis and sphingolipid degradation increased membrane rigidity. This makes it unlikely that the observed rigidification is caused by accumulation or depletion of a single shared lipid species. Rather, perturbations at different steps of sphingolipid metabolism may lead to distinct lipidomic changes that nevertheless converge on a common biophysical outcome: reduced endolysosomal membrane fluidity. In parallel, we used C-Laurdan imaging to quantify membrane order and found that SPHK2 knockdown in SH-SY5Y human neuroblastoma cells increased GP values, consistent with increased membrane rigidity. Two-channel thresholding of LysoTracker-positive compartments further showed that SPHK2 knockdown increased GP values in lysosome-associated regions.

      Thus, although lipidomics will be valuable to define the underlying lipid changes in future work, the current data already provide convergent quantitative evidence from independent membrane-fluidity readouts across C. elegans and human cell models. This cross-model consistency strengthens the robustness and reproducibility of the central conclusion that perturbation of sphingolipid metabolism alters endolysosomal membrane properties and promotes membrane rupture.

      The conclusions of the study rely heavily on imaging-based assays, including FRAP, C-Laurdan, and fluorescence microscopy. While these approaches provide valuable spatial and qualitative insights, they are inherently indirect and subject to interpretive limitations. To strengthen the mechanistic claims, the authors should incorporate additional biochemical or quantitative approaches. For example, lipidomics would allow direct measurement of membrane lipid composition changes, and western blotting or quantitative proteomics could assess levels of membrane-associated proteins involved in endolysosomal function or stress responses. Including such data would significantly improve the robustness and reproducibility of the study's conclusions.

      We agree that lipidomic and proteomic analyses will be important in future studies to define which sphingolipid species and/or membrane-associated proteins contribute to the observed rigidification of endolysosomal membranes. In response to this point, we have revised the wording throughout the manuscript to more precisely distinguish our experimental interventions from inferred changes in lipid composition. Because we did not directly measure lipid composition in the present study, we now refer more specifically to “genetic perturbation/disruption of sphingolipid metabolism” or “knockdown of enzymes involved in sphingolipid metabolism” when describing our data, rather than implying that global sphingolipid homeostasis was directly quantified. We retain “sphingolipid imbalance” only in interpretive or model-based statements where appropriate.

      However, we respectfully disagree that the current data are only qualitative. FRAP and C-Laurdan GP imaging are established quantitative biophysical approaches: FRAP provides quantitative parameters such as the time to half-maximal recovery and the mobile fraction, whereas C-Laurdan GP provides a ratiometric measurement of membrane lipid order and packing. Similarly, the Galectin puncta assay is an established quantitative readout of lysosomal membrane permeabilization. Thus, while these approaches are imaging-based, they provide quantitative readouts of membrane mobility, membrane order, and membrane rupture, respectively.

      We also note that lipidomic and proteomic profiling, although valuable, would not by itself establish which lipid or protein changes causally drive the membrane rigidification observed in our study. Membrane fluidity is an emergent biophysical property determined by the combined composition of the membrane, including lipid abundance, saturation, acyl-chain length, head groups, sterol content, and membrane-associated proteins. Therefore, an increase or decrease in a given lipid or protein species cannot be directly translated into a predictable change in membrane fluidity without additional biophysical validation. This point is further supported by our observation that knockdown of genes involved in both sphingolipid biosynthesis and sphingolipid degradation increased endolysosomal membrane rigidity. These perturbations would be expected to affect lipid composition in different, possibly even opposing, ways, making it unlikely that the shared rigidification phenotype is caused by accumulation or depletion of one single lipid species. Rather, distinct lipidomic changes may converge on a common biophysical outcome: reduced endolysosomal membrane fluidity.

      We have clarified this point in the Discussion and now state that future lysosome-enriched or organelle-specific lipidomic/proteomic approaches should be combined with direct manipulation of candidate lipid or protein species, followed by measurements of membrane fluidity and rupture, to determine which changes causally contribute to endolysosomal membrane rigidification. Such experiments would address the distinct question of which molecular components mediate the effect. By contrast, the central aim of the present study was to test whether genetic perturbation of enzymes involved in sphingolipid metabolism alters membrane fluidity and thereby promotes endolysosomal rupture and tau seeding.

      For this question, direct biophysical measurements of membrane fluidity and quantitative readouts of membrane rupture are the most relevant assays. We therefore used complementary quantitative approaches in two distinct model systems: FRAP of the lysosomal membrane protein LAAT-1::mCherry in C. elegans and C-Laurdan GP imaging in human cells. The fact that perturbing sphingolipid metabolism reduced endolysosomal membrane fluidity in C. elegans and increased lysosome-associated membrane rigidity in human cells supports the robustness and reproducibility of the central conclusion across independent model systems. In the revised manuscript, we further strengthened the human-cell data by adding SH-SY5Y neuroblastoma cells as a neuronal-like model and by combining C-Laurdan imaging with LysoTracker-based analysis to assess lysosome-associated membrane properties.

      To further address causality, we manipulated membrane fluidity independently of sphingolipid metabolism enzymes using fatty acid supplementation. Increasing membrane rigidity with PA exacerbated tau-induced endolysosomal rupture and seeded aggregation, whereas increasing membrane fluidity with ALA reduced tau-induced membrane rigidification, endolysosomal rupture, and seeded aggregation. Thus, the revised manuscript combines genetic perturbation of sphingolipid metabolism, quantitative membrane-fluidity measurements, whole-cell and lysosome-associated C-Laurdan analysis, and Galectin-based rupture assays across complementary models.

      Regarding lysosomal function, we agree that functional readouts are informative, but lysosomal membrane rupture and global lysosomal degradative capacity are related but not identical readouts. This distinction is supported by Yong et al., who reported that lipid dysregulation can induce lysosomal membrane permeabilization and lysosomal accumulation of endogenous protein aggregates without broadly impairing core lysosomal or proteasomal functions [2]. Accordingly, the absence of overt defects in general lysosomal activity would not necessarily exclude membrane damage.

      The human cell experiments were performed exclusively in HEK293T cells, which are not physiologically relevant for modeling Alzheimer's disease or lysosomal function in neurons. Given that the study aims to draw conclusions related to tau aggregation and lysosomal membrane integrity, the use of a more disease relevant cellular model is essential. There are several established AD-relevant cell models, including iPSCderived neurons, neuroblastoma lines expressing tau, or microglial models, which would better reflect the cellular context of tauopathies. Validation of key findings in at least one of these systems would substantially enhance the biological relevance and translational impact of the study.

      We have expanded and clarified the human cell data in the revised manuscript. Specifically, we now include SH-SY5Y human neuroblastoma cells for key C-Laurdan experiments assessing membrane rigidity after SPHK2 knockdown. We also show that recombinant tau fibrils increased membrane rigidity in SHSY5Y and HEK293T cells, including in LysoTracker-positive compartments.

      Importantly, the HEK293T cells are used for specific, established quantitative assays rather than as a model of neuronal toxicity. In particular, HEK293T sfGFP-LGALS3 cells are used to quantify galectin puncta formation as a readout of endolysosomal rupture, and HEK tau-Venus biosensor cells are used to quantify seeded tau aggregation. Thus, SH-SY5Y cells and HEK293T cells are used for complementary purposes: SHSY5Y cells provide a more neuronal-like human cell context for membrane-rigidity measurements, whereas HEK293T reporter/biosensor cells provide robust quantitative assays for galectin puncta formation and tau seeding.

      In addition, tau-associated neuronal dysfunction and toxicity were assessed in vivo, in functional C. elegans touch receptor neurons. In the revised manuscript, we show that ALA supplementation mitigated the age-dependent touch-response deficit and reduced neurotoxicity in animals expressing F3ΔK281::mCherry in touch receptor neurons. We have also revised the wording throughout the manuscript to avoid implying that HEK293T cells are used to model neuronal toxicity.

      Finally, the relevance of these hits to human neuronal tau seeding is also supported by our previous study, in which conserved hits from the C. elegans screen, including sphingosine kinase perturbation, were validated in human iPSC-derived neurons for their effect on seeded tau aggregation [1]. We now cite this published study where appropriate. Together, the revised manuscript combines neuronal-like human SH-SY5Y cells, established HEK293T tau-seeding and galectin reporter assays, in vivo neuronal readouts in C. elegans, and prior validation in human iPSC-derived neurons, thereby strengthening the biological relevance of the conclusions while using each model for the assay in which it is most informative.

      The authors reported that PUFA supplementation rescues neurotoxic phenotypes by increasing membrane fluidity. However, the data supporting this claim rely entirely on confocal imaging, shown in both the main and supplemental figures. To substantiate the mechanistic link between PUFA treatment and improved lysosomal membrane properties, the authors should include functional assays demonstrating that PUFAs are indeed incorporated into lysosomal membranes. Additionally, lipidomics analysis would be valuable to identify which lipid species are altered upon supplementation and correlate these changes with the observed phenotypic rescue. Furthermore, the conclusion that PUFAs rescue "neurotoxic phenotypes" is not appropriate based on data derived solely from HEK293T cells, which are not neuronal. To make claims about tau-related neurotoxicity, the authors should validate their findings in a more relevant neuronal model, such as SH-SY5Y neuroblastoma cells expressing tau or iPSC-derived neurons. This would better reflect the cellular environment of Alzheimer's disease and provide stronger support for the proposed therapeutic potential of PUFA supplementation.

      We agree that PUFA supplementation can have effects beyond membrane fluidity and that our data do not directly demonstrate incorporation of ALA into lysosomal membranes. We have therefore revised the Discussion to explicitly acknowledge this limitation. In the revised text, we state that “PUFAs can also influence lipid signaling, oxidative stress responses, and broader membrane remodeling” and that we “cannot exclude additional direct or indirect effects of ALA.” At the same time, we note that the opposing effects of PA and ALA, together with the sphingolipid-metabolism knockdown data, support membrane fluidity as a major determinant of endolysosomal membrane integrity and rupture in our models. To strengthen the link between ALA and lysosome-associated membrane properties, we combined CLaurdan imaging with LysoTracker-based analysis. In the revised Results, we show that ALA prevented tau-induced membrane rigidification and that LysoTracker-based analysis indicated effects on lysosome-associated membrane properties. ALA also reduced tau-induced endolysosomal rupture and seeded aggregation in human cell models.

      Regarding lipidomics, we refer to our response above and to the revised Discussion. We agree that lipidomics would be valuable to identify ALA-induced lipid changes, but such data would not by itself establish how these changes affect membrane fluidity without additional biophysical validation.

      Finally, we clarify that our conclusion regarding tau-associated neuronal dysfunction and toxicity is not based on HEK293T cells. HEK293T cells were used for established quantitative assays of Galectin puncta formation and seeded tau aggregation. The neurotoxicity experiments were performed in vivo in C. elegans touch receptor neurons, where ALA supplementation reduced galectin foci formation, mitigated age-dependent touch-response deficit and reduced neuronal toxicity.

      While the authors demonstrate that ALA supplementation mitigates neurotoxicity in C. elegans expressing aggregated tau (F3ΔK281::mCherry), the current data are not sufficient to conclude that ALA directly rescues tau aggregation toxicity via a lysosome-specific mechanism. It remains unclear how lipid composition is altered upon ALA treatment and whether these changes correlate with functional improvement of lysosomal pathways. The manuscript does not provide mechanistic insight into how ALA enhances lysosomal health or attenuates endolysosomal damage. Moreover, supplementation with PUFAs like ALA can activate a wide range of cellular processes beyond lysosomal function, including alterations in membrane fluidity, signaling cascades, and oxidative stress responses. The authors should clarify how they distinguish the lysosome-related effects from these alternative pathways. For example, did they observe specific lysosomal markers or structural improvements in lysosomes upon ALA treatment?

      Additional data or controls would be necessary to support a lysosome-specific protective mechanism and to exclude the involvement of other PUFA-responsive pathways in the observed phenotypes.

      We agree that our data do not prove that ALA acts exclusively through a lysosome-specific mechanism or that ALA is directly incorporated into lysosomal membranes. We have therefore revised the manuscript to avoid this interpretation and explicitly acknowledge alternative PUFA-responsive pathways. In the revised Discussion, we state that “PUFAs can also influence lipid signaling, oxidative stress responses, and broader membrane remodeling” and that we “cannot exclude additional direct or indirect effects of ALA.” We further conclude more cautiously that the opposing effects of PA and ALA, together with the sphingolipid metabolism perturbation data, support membrane fluidity as a major determinant of endolysosomal membrane integrity and rupture in our models.

      To strengthen the lysosome-related aspect of the mechanism, we added LysoTracker-based analysis to the C-Laurdan experiments. In the revised Results, ALA prevented tau-induced membrane rigidification, and LysoTracker-based analysis indicated that ALA also affected lysosome-associated membrane properties. ALA further reduced tau-induced Galectin puncta formation and seeded tau aggregation in human cell models. These data support an effect of ALA on lysosome-associated membrane order and rupture, while not excluding additional effects through lipid signaling, oxidative stress responses, or other PUFA-responsive pathways.

      Regarding lipid composition, we refer to the revised Discussion and our response above. We agree that lipidomics would be valuable to identify ALA-induced lipid changes, but such analyses would need to be organelle-specific and combined with biophysical validation to determine how candidate lipid changes affect membrane fluidity and rupture.

      Finally, we clarify that our conclusion regarding tau-associated neuronal dysfunction and toxicity is based on the C. elegans experiments, not on HEK293T cells. HEK293T cells were used for quantitative Galectin puncta and tau-seeding assays, whereas neuronal dysfunction and toxicity were assessed in vivo in touch receptor neurons in C. elegans. In the revised Results, we state that ALA supplementation mitigated galectin foci formation, age-dependent touch-response deficit and reduced neuronal toxicity in animals expressing F3ΔK281::mCherry in these neurons.

      Reviewer #2 (Public review):

      Tittelmeier et al. investigated the role of sphingolipid (SL) metabolism in the maintenance of endolysosomal vesicle integrity. They find that both impaired SL biosynthesis and degradation in C. elegans, decrease the fluidity of endolysosomal membranes and promote their rupture, while it has little effect on plasma membrane fluidity. Endolysosomal membrane fluidity is also negatively affected in human cells upon knockdown (KD) of a gene (SPHK2) involved in the SL degradation pathway. Aggregated forms of tau in both models (C. elegans and human cells) can also cause rigidification of the endolysosomal membrane, with SL homeostasis disruption having an additive effect, exacerbating endolysosomal rupture. Notably, KD of SPHK2 also increased the formation of tau foci, suggesting that compromised endolysosomal integrity may promote tau aggregation. These data provide a clearer understanding of how genetic manipulation of SL metabolism affects endolysosomal membranes and their rigidification in the context of tau aggregation. Supplementation of polyunsaturated fatty acids (PUFAs), which has a beneficial effect on Alzheimer's patients, improved membrane fluidity and reduced tau propagation in human cells and tau-associated neurotoxicity in C. elegans, suggesting a possible mechanism of action.

      Overall, the conclusions of this paper are supported by the data, with a few aspects requiring further clarification and elaboration.

      (1) A reference to Figure S2E-G, which shows that KD of SL biosynthesis genes do not affect the plasma membrane, is missing from the main text.

      We thank the reviewer for pointing this out. We have added the reference to the respective figures in the main text when discussing the plasma membrane FRAP experiments.

      (2) In Figure 3C, lipofectamine alone shows that it increases membrane rigidity (increased GP values), not membrane fluidity.

      We thank the reviewer for pointing out this incorrect wording. A higher GP index indicates increased membrane order/rigidity, not increased membrane fluidity. Since the revised main figure now includes the SH-SY5Y data, in which Lipofectamine alone did not significantly alter GP values, we removed the misleading statement from the Results. Importantly, all siRNA-mediated knockdown experiments were compared to matched control siRNA conditions exposed to the same transfection reagent. Thus, the effect attributed to SPHK2 KD is assessed relative to the appropriate Lipofectamine-containing control condition.

      (3) In Figure 3F, the EV cntl condition expressing F3:mCh tau should have increased LGALS3 foci compared to the mCh EV cntl according to Ref (20) and its Figure 2G (at least for Day 5 animals), which would be indicative of the tau spreading in hypodermal tissue. What C. elegans age was examined in Figure 3F? Can the authors provide evidence of the transmission of the F3:mCh tau from the touch receptor neurons to the hypodermis in the EV [similar to Figure 2C & D from Ref (20)] and compare it to the KDs? Otherwise, it seems that KD of SL genes impacts not only endolysosomal rupture but significantly affects tau accumulation/spreading as well (e.g., shown later in HEK cells, where SPHK2 KD increases the formation of tau-Venus foci).

      We thank the reviewer for raising this important point. The analysis referred to by the reviewer has now been moved to the revised C. elegans section and is presented as Figure 4A and B. The experiments were performed in the reporter strain used in our genome-wide screen In Ref (20), now published in Autophagy [1]. We clarified the purpose of the reporter strain and the relationship between tau transmission and the galectin puncta readout. In the revised manuscript, we now state: “In this strain, endolysosomal membrane damage is monitored in the hypodermis by expression of human galectin-3 fused to superfolder-GFP (sfGFP::LGALS3). The animals also express an aggregation-prone tau fragment fused to mCherry (F3ΔK281::mCherry) in touch receptor neurons, which is transmitted to the hypodermis, as described previously [20].” We further clarify that sfGFP::LGALS3 puncta formation, not F3ΔK281::mCherry, is the readout of endolysosomal rupture: “Under steady-state conditions, sfGFP::LGALS3 remains diffusely distributed throughout the cytosol. Upon endolysosomal damage, luminal β-galactosides become exposed and recruit sfGFP::LGALS3 into visible puncta, providing a sensitive readout of vesicle rupture”.

      Furthermore, we now better explain that transmitted tau sensitizes endolysosomal membranes to additional perturbations rather than necessarily inducing a strong lysosomal rupture phenotype on its own. In the experiments shown in Figure 4A and B, we compare F3ΔK281::mCherry animals with matched mCherry-only control animals that also express sfGFP::LGALS3 in the hypodermis. We now state: “We compared animals expressing F3ΔK281::mCherry in touch receptor neurons, from where it is transmitted to the hypodermis, with matched controls expressing mCherry alone in the same neurons. In both strains, sfGFP::LGALS3 is expressed in the hypodermis to monitor endolysosomal membrane damage.” We have also clarified the age of the animals in the revised figure legends.

      To experimentally address whether the enhanced rupture phenotype could be explained by altered tau transmission, we quantified hypodermal F3ΔK281::mCherry levels after sphk-1 RNAi (new Figure 4C, D). Importantly, sphk-1 RNAi did not increase hypodermal F3ΔK281::mCherry levels, arguing that the enhanced rupture phenotype is not due to increased tau transmission. Moreover, C. elegans neurons are largely refractory to systemic RNAi under the conditions used here [3]. We have added this important information to the Discussion. Specifically, the revised manuscript states that “the enhanced rupture phenotype is unlikely to result from a direct effect of RNAi on neuronal F3ΔK281::mCherry expression, as C. elegans neurons are largely refractory to systemic RNAi under the conditions used here,” supporting the interpretation that the RNAi treatments primarily affect endolysosomal integrity in the recipient tissue rather than neuronal tau expression itself.

      Finally, we would like to clarify that the increased tau-Venus foci in HEK cells should not be interpreted as a direct induction of tau aggregation by SPHK2 KD. Only upon addition of recombinant tau fibrils did SPHK2 KD significantly increase tau-Venus foci formation (Figure 3 H, I). This is consistent with the control experiments performed in human iPSCs in our previous study and supports our interpretation that perturbation of sphingolipid metabolism increases susceptibility to seeded tau aggregation by promoting endolysosomal rupture and tau seed escape, rather than by directly increasing tau aggregation or tau transmission.

      (4) Sphingolipids are essential membrane components and signaling molecules. Does KD of SL genes in C. elegans and the subsequent endolysosomal rupture cause any major, intermediate, or minor defects/phenotypes (in non-aggregation prone models, w/t.)?

      We agree that sphingolipids are essential membrane components and signaling molecules and that perturbing sphingolipid metabolism can have broader physiological consequences. In the revised manuscript, we address this point in two ways.

      First, we directly tested whether SL gene knockdown can induce endolysosomal rupture independently of aggregation-prone tau by using matched control animals expressing mCherry alone in touch receptor neurons while also expressing sfGFP::LGALS3 in the hypodermis. In these animals, knockdown of most SL-related hits resulted in nearly all animals displaying hypodermal sfGFP::LGALS3 foci, indicating that perturbation of SL metabolism can compromise endolysosomal integrity in the absence of transmitted F3ΔK281::mCherry (Figure 4A, B).

      Second, we have added a Discussion paragraph to place these findings into a broader physiological context. We now clarify that endolysosomal membrane rupture and global lysosomal function are related but not identical readouts. In support of this distinction, we discuss work showing that lipid dysregulation can induce lysosomal membrane permeabilization and lysosomal accumulation of endogenous protein aggregates without broadly impairing core lysosomal or proteasomal function [2]. Thus, membrane damage can occur even when general lysosomal activity is not overtly disrupted.

      We also discuss a recent study published during the revision of this manuscript that independently identified SPHK-1 as an important regulator of lysosomal integrity in C. elegans, showing that strong sphk1 loss-of-function causes lysosomal sphingosine accumulation, membrane rupture, impaired degradative function, cargo accumulation, developmental defects, and reduced lifespan [4].

      Importantly, while that study focused on a strong loss-of-function mutation in a single SL-metabolism gene, our data show that knockdown of multiple SL-metabolism genes, including genes involved in both SL biosynthesis and degradation, converges on reduced endolysosomal membrane fluidity and increased rupture. This suggests that the observed membrane rigidification and rupture are not specific to one mutant background but represent a broader consequence of perturbing SL metabolism at multiple points. A systematic characterization of all organismal phenotypes caused by each SL gene knockdown was beyond the scope of the present study. Therefore, the revised manuscript now makes clear that the study focuses on endolysosomal membrane fluidity and rupture because these membrane-level changes are directly linked to tau seed escape and seeded tau aggregation, while broader physiological consequences may vary depending on the strength and context of the perturbation.

      Reviewer #3 (Public review):

      Summary:

      The authors set off with an analysis of the lysosomal integrity upon knockdown of genes of the sphingolipid metabolic pathway that they identified in a previous (yet unpublished) work of an RNA screen using a new C. elegans Tau model. They then used cell culture and C. elegans experiments to study the link between lysosomal rupture and Tau propagation.

      Strengths:

      The authors use two complementary model systems and use probes to assess membrane rigidity that allow a quick assessment of the membrane dynamics and offer the opportunity to treat the cells with lipids, RNAi. Tau seeds, etc.

      Weaknesses:

      The main weakness is that this work builds on not-yet-peer-reviewed manuscript that established a new C. elegans Tau model and RNAi screen that aimed to identify genes involved in the propagation of Tau.

      This reviewer misses essential information of the C. elegans Tau strain (not included in the method section): e.g., promoter used for the expression, information on the used Tau variant, expression pattern, and aggregation, etc.

      We thank the reviewer for raising this point. The related study establishing the C. elegans tau transmission model and RNAi screen has now been peer-reviewed and published in Autophagy [1]. We now cite the published article throughout the revised manuscript instead of the previous preprint.

      We also agree that the current manuscript should be understandable without requiring the reader to consult the previous paper for the basic logic of the model. We therefore added a clearer introduction of the reporter strain in the Results. Specifically, we now explain that the strain expresses the aggregation-prone tau fragment F3ΔK281::mCherry in touch receptor neurons, that this tau fragment is transmitted to the hypodermis, and that endolysosomal membrane damage is monitored in the hypodermis using sfGFP::LGALS3. We further clarify that sfGFP::LGALS3 remains diffuse under steady-state conditions and forms puncta upon endolysosomal membrane damage, when luminal β-galactosides become exposed. Thus, the revised manuscript now provides the key information needed to understand the experimental system used here, while the published Autophagy paper is cited for the full characterization of the tau transmission model, expression pattern, aggregation properties, and original genome-wide RNAi screen.

      Throughout the study, I missed data on:

      (1) Effect of the knockdown on Tau expression, localisation (with lysosomal membrane?), aggregation, and proteotoxicity. The effect of the RNAi-mediated knockdown could also simply lead to a reduced expression of Tau that, in turn, leads to suppressed propagation.

      We agree that it is important to distinguish effects on tau expression/transmission from effects on endolysosomal membrane integrity. In the C. elegans experiments, F3ΔK281::mCherry is expressed in touch receptor neurons and transmitted to the hypodermis, where sfGFP::LGALS3 reports endolysosomal membrane damage. We now describe this more clearly in the revised Results.

      The RNAi treatments target genes involved in sphingolipid metabolism under systemic RNAi conditions. Because C. elegans neurons are largely refractory to systemic RNAi in the absence of sensitizing backgrounds [3], which we did not use, a direct RNAi-mediated reduction of neuronal F3ΔK281::mCherry expression is very unlikely. We have added this point to the Discussion, stating that “the enhanced rupture phenotype is unlikely to result from a direct effect of RNAi on neuronal F3ΔK281::mCherry expression, as C. elegans neurons are largely refractory to systemic RNAi under the conditions used here.”

      Experimentally, we also tested whether sphingolipid perturbation alters transmitted tau levels (new Figure 4C, D). Specifically, we quantified hypodermal F3ΔK281::mCherry after sphk-1 RNAi and found no increase, arguing that the enhanced rupture phenotype is not due to increased tau transmission.

      Moreover, in the cell-based tau-Venus assay, SPHK2 knockdown alone did not induce detectable tau aggregation in the absence of exogenously added tau fibrils (Figure 3H, I). Only upon addition of recombinant tau fibrils did SPHK2 knockdown significantly increase tau-Venus foci formation. We now state this explicitly in the revised Results and conclude that disruption of sphingolipid metabolism is not sufficient on its own to initiate detectable tau aggregation under the conditions tested here but rather increases cellular susceptibility to seeded tau aggregation when tau fibrils are present.

      Together, these data argue against a direct effect of sphingolipid gene knockdown on tau expression or spontaneous tau aggregation. Instead, they support our interpretation that perturbation of sphingolipid metabolism compromises endolysosomal membrane integrity, thereby facilitating tau seed escape and seeded aggregation when tau seeds are present.

      (2) A quantification of RNAi knockdown is needed to judge the efficiency of the RNAi, in particular for the combinatorial RNAi experiments involving 2 and even 4 genes in parallel. Ideally, these analyses should be validated with mutants for these genes.

      We agree that RNAi efficiency can vary between clones and that this is particularly relevant for combinatorial RNAi experiments targeting two or more potentially redundant genes. We have added this limitation to the Results section and now state: “Because KD efficiency was not assessed for the individual RNAi clones or co-RNAi combinations, these experiments do not allow comparison of relative RNAi strength or inference of the relative importance of individual genes. Thus, the conclusions drawn from these RNAi experiments are qualitative: specific single or combined KDs can promote endolysosomal rupture, whereas the absence of a detectable phenotype after RNAi cannot exclude gene involvement, as KD may have been insufficient.”

      Where mutant strains were available, we performed genetic validation. Specifically, a sphk-1 mutant available at CGC (CZ24969; sphk-1(ju831)) also showed increased hypodermal sfGFP::LGALS3 puncta (new Figure S1A), supporting the RNAi-based conclusion that genetic perturbation of sphingolipid metabolism compromises endolysosomal integrity. Corresponding mutant strains were not available for the other selected hits. Importantly, most hits also induced sfGFP::LGALS3 foci in human HEK293T cells as assessed in our previous study [1], providing additional support that the observed effects are not random RNAi artifacts.

      Further:

      (3) Figure 4 H, I: Would Tau also aggregate in the absence of externally added Tau?

      No. In the tau-Venus biosensor cell line, SPHK2 knockdown alone did not increase tau-Venus foci formation (now Figure 3H, I). Tau-Venus foci increased only after addition of recombinant tau fibrils and were further enhanced by SPHK2 knockdown. We now state this explicitly in the Results.

      (4) How specific is the effect for Tau? It would help if the authors could assess other amyloid proteins.

      We agree that similar membrane-level mechanisms may apply to other amyloid assemblies. We have therefore added recent literature to the Discussion supporting the broader concept that intralysosomal amyloid assemblies can physically deform and rupture lysosomal membranes. The revised manuscript states: “This interpretation is consistent with recent ultrastructural studies showing that intralysosomal amyloid assemblies can physically deform and rupture lysosomal membranes.” We further clarify that “whether this mechanism is specific to tau or also applies to other amyloid assemblies remains to be determined.”

      Whether perturbation of SL metabolism similarly affects endolysosomal escape and seeded aggregation of other disease-associated amyloid proteins is an important question that we plan to address in future work. However, these experiments require additional disease-specific models, aggregation assays, and validation, and are therefore beyond the scope of the present revision.

      (5) The connection between sphingolipids and AD is not new. See He et al, 2010, Neurobiol. Aging + numerous publications and also not between Tau seeding and lysosomal rupture: Rose et al., PNAS 2024 (that has been cited by the authors).

      We agree and our manuscript does not aim to establish these associations as new. We state explicitly that alterations in sphingolipid metabolism have been reported in aging and AD, and that endolysosomal rupture is increasingly recognized as a critical step in tau seed escape and propagation.

      The novelty of our study lies in mechanistically connecting these two previously established areas. Specifically, we show that genetic perturbation of enzymes involved in sphingolipid metabolism reduces endolysosomal membrane fluidity, promotes membrane rupture, and thereby increases susceptibility to tau seed escape and seeded aggregation. We have revised the Introduction and Discussion to better emphasize this mechanistic contribution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Figure formatting and annotation need improvement. Panel letters throughout the figures should be in uppercase, and gene names in pathway diagrams should be italicized for consistency. Several scale bars are missing, including in Figures 1C, 2A, and 2H, and should be clearly indicated in the figures and legends. In Figure 1C, the age of the worms used in the assay is not specified. While the Methods section mentions "age-synchronized animals," the precise age at the time of imaging or experimentation is not stated. It would strengthen the study to explore whether membrane integrity phenotypes vary between young adults (day 1) and older adults (day 7 or 10) across the different conditions. Figure 1B lacks sufficient detail describing the galectin puncta assay used. A brief explanation of the assay rationale and readout would help contextualize the findings.

      We thank the reviewer for pointing this out. We have revised the figures and figure legends accordingly by standardizing panel labels, adding scale bars where missing, and providing the age of animals used in the assays. We also expanded the description of the galectin puncta assay in the Results to explain the rationale and readout of sfGFP::LGALS3 puncta formation.

      Regarding the reviewer’s suggestion to compare young and aged animals, we agree that age-dependent changes in endolysosomal membrane integrity are an interesting question. However, the purpose of the present study was to investigate how perturbation of sphingolipid metabolism affects endolysosomal membrane fluidity and rupture under the assay conditions used in our original screen. A systematic comparison across aging is beyond the scope of the current revision. We have therefore clarified the animal ages used in the relevant figure legends and Methods.

      In Figure S1A, the authors show co-knockdown of multiple genes, including one condition with simultaneous RNAi against four targets. Because different RNAi clones can vary in knockdown efficiency, it is important to provide validation of gene knockdown levels (e.g., by qRT-PCR) shown in both panels a and b.

      We agree that RNAi efficiency can vary between clones and that this is particularly relevant for combinatorial RNAi experiments targeting two or more potentially redundant genes. We have added this limitation to the Results section and now state: “Because KD efficiency was not assessed for the individual RNAi clones or co-RNAi combinations, these experiments do not allow comparison of relative RNAi strength or inference of the relative importance of individual genes. Thus, the conclusions drawn from these RNAi experiments are qualitative: specific single or combined KDs can promote endolysosomal rupture, whereas the absence of a detectable phenotype after RNAi cannot exclude gene involvement, as KD may have been insufficient.”

      Where mutant strains were available, we performed genetic validation. Specifically, a sphk-1 mutant available at CGC (CZ24969; sphk-1(ju831)) also showed increased hypodermal sfGFP::LGALS3 puncta (new Figure S1A), supporting the RNAi-based conclusion that genetic perturbation of sphingolipid metabolism compromises endolysosomal integrity. Corresponding mutant strains were not available for the other selected hits. Importantly, most hits also induced sfGFP::LGALS3 foci in human HEK293T cells as assessed in our previous study [1], providing additional support that the observed effects are not random RNAi artifacts.

      In Figure 2E, the FRAP recovery curves show only ~60% recovery in controls after 25 seconds, and an even lower recovery (~40%) in hpo-8 and spp-10 RNAi conditions. The authors should discuss why the recovery is incomplete and what it implies about the mobile fraction of the protein or membrane components in these conditions.

      We agree that incomplete FRAP recovery is informative. For this reason, we report both the time to half-maximal recovery (thalf) and the maximal recoverable fluorescence signal. Increased thalf indicates reduced lateral mobility of LAAT-1::mCherry within the lysosomal membrane, consistent with reduced membrane fluidity. In addition, a reduced maximal recovery suggests that a larger fraction of the reporter is immobile or only slowly mobile during the time window analyzed. This may reflect stronger confinement of LAAT-1::mCherry within even more rigid membrane domains. However, because RNAi efficiency may differ between clones and we have not assessed their individual KD efficiency, we avoid overinterpreting differences in the absolute strength of recovery defects between individual KDs. Instead, we conclude that KD of sphingolipid-metabolism genes identified in our screen consistently reduces lysosomal membrane fluidity, as reflected by increased thalf and, in some cases, reduced maximal recovery.

      In Figure S3A, the Western blot for SPHK2 shows unequal loading between the control and siSPHK2 lanes. The blot should be normalized to a loading control and quantified to demonstrate knockdown efficiency.

      We have quantified SPHK2 levels relative to GAPDH across independent experiments and present the normalized quantification (Figure S3C-E).

      Key experimental details are missing from the manuscript. The strains of C. elegans and RNAi bacteria used were not described, and there is no information on biological replicates. The authors should clarify how many times each experiment was performed and provide more transparency on experimental reproducibility.

      We thank the reviewer for pointing this out. The C. elegans strains and RNAi bacterial clones used in this study were established and fully described in our previous study, which has now been published in Autophagy [1]. We now cite the published article throughout the revised manuscript and have added additional information in the Results section to explain the key features of the strains used here.

      We have also revised the Methods and figure legends to improve transparency regarding experimental details. The figure legends include the number of biological replicates, the number of animals or cells analyzed, and the statistical tests used for each experiment. In addition, the Statistical Analysis section in the Methods now summarizes how replicate numbers and sample sizes are reported across the study. Finally, the source details for the strains and RNAi clones used in this study are now provided in Tables S1 and S2, respectively. These revisions should improve the experimental clarity and reproducibility of the data shown.

      References:

      (1) Sandhof CA, Martin N, Tittelmeier J, Schlueter A, Pezzali M, Schoendorf DC, et al. A novel C. elegans model for MAPT/Tau spreading reveals genes critical for endolysosomal integrity and seeded MAPT/Tau aggregation. Autophagy. 2025;21(12):2963-81. Epub 20250904. doi: 10.1080/15548627.2025.2551676. PubMed PMID: 40851193; PubMed Central PMCID: PMCPMC12758218.

      (2) Yong J, Villalta JE, Vu N, Kukurugya MA, Olsson N, Lopez MP, et al. Impairment of lipid homeostasis causes lysosomal accumulation of endogenous protein aggregates through ESCRT disruption. eLife. 2024;12. Epub 20241223. doi: 10.7554/eLife.86194. PubMed PMID: 39713930; PubMed Central PMCID: PMCPMC11666243.

      (3) Calixto A, Chelur D, Topalidou I, Chen X, Chalfie M. Enhanced neuronal RNAi in C. elegans using SID-1. Nat Methods. 2010;7(7):554-9. doi: 10.1038/nmeth.1463. PubMed PMID: 20512143; PubMed Central PMCID: PMC2894993.

      (4) Li Y, Zhang J, Li M, Yang L, Wang X. Sphingosine kinase SPHK-1 maintains sphingolipid metabolism to protect lysosome membrane integrity in C. elegans. Mol Biol Cell. 2026;37(1):ar1. Epub 20251105. doi: 10.1091/mbc.E25-04-0182. PubMed PMID: 41191545; PubMed Central PMCID: PMCPMC12696880.

    1. eLife Assessment

      This study presents a fundamental finding that the JAK-STAT pathway (JSP) exerts context-dependent roles across distinct cellular compartments within the breast cancer microenvironment. The conclusions are supported by convincing evidence from multi-omics analyses. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

    2. Reviewer #2 (Public review):

      Summary:

      The JAK-STAT pathway (JSP) exhibits cell-type-specific functional heterogeneity in breast cancer. This study investigates the JSP in breast cancer and its response to anti-PD‑1 immunotherapy. JSP displays distinct cell‑type heterogeneity: it promotes malignant phenotypes and immunosuppression in tumor cells, while enhancing cytotoxicity and reducing exhaustion in T cells. Elevated JSP expression correlates with improved immunotherapy responses, especially in triple‑negative breast cancer. These findings highlight the paradoxical roles of JSP, indicating that broad inhibition may compromise anti‑tumor immunity.

      Strengths:

      The major strengths of this study include the comprehensive characterization JSP heterogeneity across epithelial, tumor, and T cells in breast cancer. The identification of JSP and STAT4 as predictive biomarkers for immunotherapy response, particularly in triple‑negative breast cancer, provides clinically relevant insights for patient stratification.

      Comments on revised version.

      The corresponding content has been revised.

    3. Reviewer #3 (Public review):

      Summary:

      This multi-omics study by Zhou et al elucidates the context-dependent roles of the Janus kinase-signal transducer and activator of transcription (JAK-STAT) pathway (JSP) across different cellular compartments in the breast cancer tumor microenvironment. While bulk JSP activity is associated with a favorable prognosis, single-cell analysis reveals a paradoxical landscape: high JSP in T cells drives anti-tumor cytotoxicity and reduces exhaustion, whereas high activity in tumor epithelial cells promotes malignancy and immunosuppression via the MIF-CD74 signaling axis. The JSP score (immune-related) serves as a robust predictive biomarker for response to anti-PD-1 immunotherapy, particularly in triple-negative breast cancer (TNBC). Furthermore, the study identifies the STAT4/SLC47A1 axis as a critical mechanism through which tumor cells resist ferroptosis, facilitating disease progression. These findings suggest that broad JAK-STAT inhibition may be counterproductive in cancer therapeutics; instead, therapeutic success depends on precise modulation and carefully timed interventions to preserve its T-cell-associated functions. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

      Strengths:

      Significant therapeutics implications

      Weaknesses:

      Limited molecular mechanisms

      Comments on revised version:

      The authors have addressed my comments

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In their manuscript, Zhou and colleagues present a detailed look at how the JSP functions differently in the various cells of a breast tumor. The authors have effectively shown that the JSP acts as a double-edged sword, as it helps T cells fight cancer but also allows tumor cells to grow and avoid ferroptosis. These findings are important because they identify a useful biomarker to predict how TNBC patients might respond to PD-1 inhibitors.

      Strengths:

      This work is important because it provides a clear explanation for the conflicting roles of the JSP in the tumor environment. The evidence is solid, as it combines data from thousands of patients with single-cell analysis and lab experiments to confirm the role of STAT4 in cancer progression and immunity.

      Comments on revised version:

      The authors made a significant effort to improve the manuscript. My comments were sufficiently addressed.

      We sincerely appreciate your careful review and positive feedback. We are glad to hear that you are satisfied with the revised manuscript and acknowledge the scientific value and solid evidence of our work. Thank you again for all your efforts and valuable suggestions.

      Reviewer #2 (Public review):

      Summary:

      The JAK-STAT pathway (JSP) exhibits cell-type-specific functional heterogeneity in breast cancer. This study investigates the JSP in breast cancer and its response to anti-PD‑1 immunotherapy. JSP displays distinct cell‑type heterogeneity: it promotes malignant phenotypes and immunosuppression in tumor cells, while enhancing cytotoxicity and reducing exhaustion in T cells. Elevated JSP expression correlates with improved immunotherapy responses, especially in triple‑negative breast cancer. These findings highlight the paradoxical roles of JSP, indicating that broad inhibition may compromise anti‑tumor immunity.

      Strengths:

      The major strengths of this study include the comprehensive characterization JSP heterogeneity across epithelial, tumor, and T cells in breast cancer. The identification of JSP and STAT4 as predictive biomarkers for immunotherapy response, particularly in triple‑negative breast cancer, provides clinically relevant insights for patient stratification.

      Weaknesses:

      The corresponding content has been revised.

      We sincerely thank you for your detailed review and valuable comments. We greatly appreciate your recognition of the cell-type-specific heterogeneity of the JAK-STAT pathway and the clinical value of JSP and STAT4 as predictive biomarkers for immunotherapy in triple-negative breast cancer. We have thoroughly revised the manuscript according to your previous suggestions, and all raised concerns have been fully addressed.

      Reviewer #3 (Public review):

      Summary:

      This multi-omics study by Zhou et al elucidates the context-dependent roles of the Janus kinase-signal transducer and activator of transcription (JAK-STAT) pathway (JSP) across different cellular compartments in the breast cancer tumor microenvironment. While bulk JSP activity is associated with a favorable prognosis, single-cell analysis reveals a paradoxical landscape: high JSP in T cells drives anti-tumor cytotoxicity and reduces exhaustion, whereas high activity in tumor epithelial cells promotes malignancy and immunosuppression via the MIF-CD74 signaling axis. The JSP score (immune-related) serves as a robust predictive biomarker for response to anti-PD-1 immunotherapy, particularly in triple-negative breast cancer (TNBC). Furthermore, the study identifies the STAT4/SLC47A1 axis as a critical mechanism through which tumor cells resist ferroptosis, facilitating disease progression. These findings suggest that broad JAK-STAT inhibition may be counterproductive in cancer therapeutics; instead, therapeutic success depends on precise modulation and carefully timed interventions to preserve its T-cell-associated functions. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

      Strengths:

      Significant therapeutics implications

      Weaknesses:

      Limited molecular mechanisms

      Comments on revised version:

      The authors have addressed my comments

      Many thanks for your careful evaluation and valuable suggestions. We highly appreciate your affirmation of the therapeutic significance of this study. We have fully revised the manuscript to enrich the molecular mechanisms, and all your comments have been properly resolved.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The most content has been revised.

      Minor corrections:

      (1) The icon about "prognosis" in graphic abstract is overly childish.

      The graphical abstract has been redrawn. The inappropriate prognosis icon is deleted accordingly.

      (2) Please double check the whole content to avoid typos. For instance, "2.2" and "2.3" have been repeated twice.

      We appreciate your reminder. The duplicate numbering of 2.2 and 2.3 resulted from Word’s automatic heading feature. We have disabled this function and fixed all repeated section numbers. In addition, we have carefully checked the full text and corrected all typos.

      (3) It will be more interesting if the oncogenic role of STAT4 could be verified via cell cloning assay.

      We appreciate your thoughtful comment. Considering the limited revision time, we cannot add the cell cloning assay in the current version. Our present data sufficiently validate the oncogenic function of STAT4, and the main conclusions remain reliable.

      Reviewer #3 (Recommendations for the authors):

      I recommend publishing the revised manuscript in eLife.

      We sincerely thank you for your positive evaluation and endorsement for the publication of our revised manuscript. We greatly appreciate your rigorous review and insightful comments that have substantially improved the quality and readability of this work.

      We sincerely appreciate all reviewers and editors for your thorough reviewing work and thoughtful feedback. Your suggestions have helped us greatly improve this manuscript. Thank you very much.

    1. eLife Assessment

      In this valuable study, the authors performed cell-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior (NT1) vs middle & posterior (NT2-9) cells of the C. elegans intestine, under fed, starved, or refeeding conditions. The data generated will be very helpful to the C. elegans community, and the evidence supporting the conclusions of the study is assessed to be solid. Some methodological caveats remain and are discussed.

    2. Reviewer #1 (Public review):

      Summary

      In this study, the authors have performed tissue-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior vs posterior cells of the C. elegans intestine. They have performed this analysis in fed and fasted states of the animal. The data generated will be very useful to the C. elegans community, and the role of pyruvate shown in this study will result in interesting follow-up investigations.

      However, several strong claims made in the study are solely based on in silico predictions and are not supported by experimental evidence.

      Comments on revised version.

      The authors have been responsive to the comments, but have not added new experiments in this manuscript that would have clarified and improved some of the mentioned shortcomings of the study.

      There are 3 comments that the authors should address:

      (1) In their response to reviewers, the authors agree that "the Pges-1deltaB promoter is not absolutely restricted to INT1 and that weak GFP expression can also be detected in INT2." They also mention that "because Pges-1deltaB is an engineered promoter derived from the intestine-specific Pges-1 promoter, this low-level INT2 expression is not unexpected." However, in line 93 of the revised manuscript, the authors claim that "Pges-1deltaB is strictly expressed in INT1 cells". This discrepancy should be fixed. They should instead describe this in line 93 as "Pges-1deltaB expression is very strongly enriched in INT1 cells, but low-level expression in INT2 was also detected".

      (2) In response to reviewers, the authors explained that "Our model is that fasting induces INS-7 secretion by lowering intracellular pyruvate in INT1 cells. Under this framework, blocking mitochondrial pyruvate breakdown would be expected to reduce pyruvate utilization and thus maintain intracellular pyruvate, preventing the drop in pyruvate that normally occurs during fasting. This would explain why these manipulations suppress fasting-induced INS-7 secretion." However, the effect of blocking import of pyruvate from cytosol into mitochondria (via knockdown of mitochondrial pyruvate carrier genes mpc-1 and mpc-2) does not agree with their proposed model. Blocking mitochondrial import of pyruvate should maintain cytosolic pyruvate levels and thus prevent the drop in pyruvate that normally occurs during fasting. In such a scenario, we would expect to see no increase in INS-7 secretion during fasting, which is opposite to the result in Fig.7D. If the pyruvate sensor is in the cytosol, we would expect that the mpc-1/2 RNAi treated animals would be unable to increase INS-7 secretion upon starvation. If the pyruvate sensor is in the mitochondrial matrix, we would expect that the mpc-1/2 RNAi treated animals would have higher INS-7 secretion than vector RNAi control animals in fed conditions. How do the authors explain this discrepancy between their observed results and their proposed model? Why does blocking mitochondrial import of pyruvate affect only refeeding-induced reduction in INS-7 secretion but not fasting-induced increase in INS-7 secretion? Is it possible that instead of responding to absolute intracellular concentrations of pyruvate, the pyruvate sensor increases INS-7 secretion upon detecting a relative drop in the mitochondrial levels of pyruvate (or its downstream metabolite)? This should be described in the text to better interpret the mpc-1/2 RNAi results.

      (3) Line 493: The authors refer to 'Table S4', which is not included in the manuscript.

    3. Reviewer #3 (Public review):

      In this study, Liu and colleagues utilize TRAP-seq to profile the repertoire of actively translated mRNAs in different intestinal cell types (anterior INT1 vs. posterior INT2-9 cells) in C. elegans. A key goal of this study was to identify transcripts differentially expressed/translated between these intestinal cell subtypes in the context of animals being well fed or subjected to acute (30 minutes) or chronic (3 hours) starvation, followed by refeeding.

      The authors identify a number of differentially expressed genes across all of the conditions tested. They then provide an initial survey of the landscape of translatome changes through Weighted Gene Network Correlation Analysis (WGNA), and some high-level functional surveys via Gene Ontology (GO) term analysis and protein domain analysis. The authors validate the enriched expression patterns of some of their identified candidate genes using fluorescent promoter fusion reporters, confirming INT1-specific expression. The authors further implicate the role of several other candidate genes in pathogen avoidance and in response to nutritional cues by knocking them down specifically in INT1 cells by RNAi. Finally, the authors identify pyruvate as a major nutrient signal coming from the bacterial diet that suppresses the release of a key insulin peptide (INS-7) and identify some of the genes expressed in INT1 that are required for this response.

      Strengths:

      (1) Good use of and justification for TRAP-seq, because scRNA-seq would be difficult under the varied conditions used (starvation, refeeding)

      (2) The manuscript is generally clear to read, and the data are generally well-presented with good supporting data that includes replicates, sample sizes, error measurements, and associated statistics.

      (3) The dataset will be an interesting resource to mine for future studies focusing on mechanisms of how particular intestinal cell types respond to different environmental signals.

      Weaknesses:

      (1) A limitation of TRAP-seq, although powerful, is that only relative comparisons can be made between genotypes/conditions to identify differentially-expressed genes, rather than assessing whether a given gene is expressed at a certain level in a cell type under a certain condition. This limitation is due to the non-specific association of sticky RNA species to the beads during the immunoprecipitation step. This is a minor point however, and the authors do a nice job of focusing their analysis on differentially expressed transcripts in the current study.

      (2) Another limitation of the current study is that the experiments testing the role of candidate genes identified by their profiling experiments do not dive a bit deeper into providing a mechanistic understanding of the phenotypes being studied. At present, the results are thus viewed more as a genomics-based screen with some limited follow-up on interesting hits. However, this reviewer appreciates that when placed in context of the work presented, a presentation of the profiling data along with some validation is an excellent starting point for future mechanistic studies elaborating on these interesting candidates.

      Appraisal of whether the authors achieved their aims, and whether the results support their conclusions.

      The main goal of the study was to survey the dynamic responses at the level of actively translated mRNAs of the INT1 vs INT2-9 cells in response to metabolic challenge.

      Overall, the authors use established methods to perform their genome-wide analysis, and the set of differentially regulated genes are enriched for expected molecular functions and form coherent networks in anticipated pathways.

      The validation experiments (promoter::GFP fusion reporters, INT1-specific knockdowns of highly regulated genes) further corroborate the quality of the TRAP-seq datasets generated.

      I have a few points for the authors that would further strengthen this work:

      (1) The authors rightfully focus on the top differentially-regulated candidates, but it's unclear at present how far down their fold change list would lead to expression pattern validations. It would be useful to test a few more promoter::GFP fusion reporters at different enrichment/fold-change/statistical cutoffs.

      (2) Although the INT1-specific RNAi provides a convenient strategy for rapidly perturbing and testing genes of interest for phenotypes, independently validating the knockdowns with genetic mutants, or alternatively (if genes are essential), degron alleles.

      Likely impact of the work on the field, and the utility of the methods and data to the community.

      The TRAP-seq data and list of differentially-expressed candidate genes will form an interesting set of high-priority candidates to study for their role in the reception and transduction of nutritional cues in response to food status and pathogens. This data will thus benefit the C. elegans community of researchers studying the mechanisms governing these phenomena.

      Comments on revised version:

      I think the authors have done a good job of addressing the suggestions from the previous round of review in this new version.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      In this study, the authors have performed tissue-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior vs posterior cells of the C. elegans intestine. They have performed this analysis in fed and fasted states of the animal. The data generated will be very useful to the C. elegans community, and the role of pyruvate shown in this study will result in interesting follow-up investigations.

      However, several strong claims made in the study are solely based on in silico predictions and are not supported by experimental evidence.

      Strengths:

      Several studies in the past have predicted different functions of the anterior (INT1) vs posterior (INT2-9) epithelial cells of the C. elegans intestine based on their anatomy and ultrastructure, but detailed characterization of differences in gene expression between these cell types (and whether indeed these are different 'cell types') was lacking prior to this study. The genes and drivers identified to be exclusively expressed in the anterior vs posterior segments of the intestine will be very helpful to selectively modulate different parts of the C. elegans intestine in future studies.

      Another strength of this study is the careful experimental design to test how the anterior vs posterior cell types of the intestine respond differently to food deprivation and recovery after return to food. These comparisons between 'states' of a cell in different physiological conditions are difficult to pick up in single-cell analyses due to low sequencing depth, which can fail to identify subtle modulation of gene expression.

      The TRAP-associated bulk RNA-seq approach used in this study is more suitable for such comparisons and provides additional information on post-transcriptional regulation during metabolic stress.

      A key finding of this study is that pyruvate levels modulate the translation state of anterior intestinal cells during fasting. Characterization of pyruvate metabolism genes, especially of the enzymes involved in its mitochondrial breakdown, provides novel insights into how gut epithelial cells respond to the acute absence of food.

      Weaknesses:

      Unlike previous TRAP-seq studies (PMID: 30580965, 36044259, 36977417) that reported sequencing data for both input and IP samples, this study only reports the sequencing data for IP samples. Since biochemical pulldowns are variable across replicates, it is difficult to know if the observed differences between different conditions are due to biological factors or differences in IP efficiency. More importantly, since two different TRAP lines were utilized in this study and a large proportion of the results focus on the differences between the translational profiles of INT1 vs INT2-9 cells, it is essential to know if the IP worked with similar efficiency for both TRAP strains that likely have different expression levels of the HA-tagged ribosomal protein. One way to estimate this would be to perform qRT-PCR of genes that are known to be enriched in all intestinal cells and determine whether their fold-enrichment over housekeeping genes (normalized to input) is similar in INT1 vs INT2-9 TRAP strains and across the fed vs fasted conditions. The authors, in fact, mention variability across biological replicates, due to which certain replicates were excluded from their WGCNA analysis.

      We appreciate the reviewer's comments. We agree that the lack of matched input sequencing libraries limits our ability to directly assess IP efficiency across replicates, conditions, and TRAP strains. However, several features of the dataset support the conclusion that the major differences reported here reflect biological rather than purely technical variation. First, the RPL-22-3xHA construct was integrated into each line to improve consistency across experiments. Second, although the INT2-9 TRAP strain yielded more RNA than the INT1 strain, as expected given the larger number of labeled cells, downstream analyses were performed on normalized count data rather than raw counts. Third, principal component analysis showed robust separation by promoter identity across all conditions, and expected INT1-enriched genes such as ins-7 were recovered in the INT1 dataset. Together, these observations support the interpretation that the TRAP datasets capture reproducible, cell-type-specific differences in ribosome-associated transcripts. Nonetheless, we agree that direct input-normalized measurements would further strengthen the study, and we will explicitly note this as an important limitation.

      It appears that GFP expression is also detectable in INT2 (in addition to strong expression in INT1 in Fig.1A). Compared to INT3-9, which looks red, INT2 cells appear yellow, suggesting that the expression patterns of the two TRAP drivers are not mutually exclusive, which changes the interpretation of many of the results described in the study.

      We agree that the Pges-1ΔB promoter is not absolutely restricted to INT1 and that weak GFP expression can also be detected in INT2. Because Pges-1ΔB is an engineered promoter derived from the intestine-specific Pges-11 promoter, this low-level INT2 expression is not unexpected. However, we note that the expression level in INT1 is substantially higher than in INT2. Thus, although the expression patterns of the two TRAP drivers are not completely mutually exclusive, Pges-1ΔB still provides the most selective available tool for enriching the INT1 translatome in the context of the current study.

      Some parts of the study overemphasize the differences between the INT1 vs INT2-9 cell types, which is a biased representation of the results. For example, the authors specifically point out that 270 genes are differentially expressed in opposite directions in INT1 vs INT2-9 cell types during acute (30 min) fasting without mentioning the 1,268 genes that are differentially expressed in the same direction. They also do not mention here that 96% of the genes are differentially expressed in the same direction in INT1 and INT2-9 cell types after prolonged (180 min) fasting, suggesting that the divergent translational responses of these cell types are only observed in the first 30 minutes of food deprivation. Similar results have also been reported for the effect of fasting on locomotory and feeding behaviors, where 30 min of fasting produces more variable effects, which become more consistent after longer periods of fasting (PMID: 36083280). Hence, the effects of brief food deprivation should be interpreted with caution.

      The intestine functions as a discrete and cohesive organ, so the expected result is that there would be no differences across the different cell types. For us, the surprise was that, in fact, there are differences between these cells at all. However, the point is well taken, and we have added a statement in the text to reflect that many genes change similarly in INT1 and INT2-9, while the differences reflect important functional divergence between these cell types.

      Many of the interpretations of this study primarily rely on pathway enrichment analyses, which are based on the known function of genes. The function of uncharacterized genes that were found to be differentially expressed in INT1 vs INT2-9 cell types, e.g., the ShKT proteins, was not explored in this study. In addition, overreliance on pathway enrichment tools (instead of functional validation) has resulted in several conflicting findings. For example, one of the main messages of this study is that INT1 cells specialize in immune and stress response in response to fasting, which relies on pathway analysis in Figs 5E and 5F. However, pathway analysis at a different time point (shown in Figure S5A) indicates that INT2-9 cells show a much stronger increase in translation of stress and pathogen-responsive genes compared to INT1 cells. Hence, some of the results should be interpreted as different translational effects in INT1 vs INT2-9 cells after different lengths of food deprivation, without making broad claims about selective pathways being affected only in specific cell types.

      We agree that some interpretations in the manuscript relied heavily on pathway enrichment analyses and should be stated more cautiously. In particular, we agree that the current data are most consistent with state-dependent differences in translational responses between INT1 and INT2-9 cells across different durations of food deprivation, rather than with the strongest version of a claim that specific pathways are selectively engaged only in one intestinal subset. We also agree that uncharacterized genes, including the ShKT family, were not mechanistically explored in the present study and should be presented as important candidates for future investigation.

      The authors have compared their TRAP-seq results with genes enriched in the anterior and posterior intestine clusters from a previously published whole-animal adult scRNA dataset (PMID: 37352352). They claim that their TRAP-seq results are in agreement with the findings of the scRNA study. However, among the 10 genes from the 'posterior intestine' scRNA cluster in Fig.S1E, six are downregulated in the INT1 vs INT2-9 comparison, while four are upregulated. Hence, there is no clear agreement between the two studies in terms of the top enriched genes in the anterior vs posterior intestine, which should be considered for cross-study comparisons in the future.

      We have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. We instead assessed the expression levels of our INT1 up-regulated genes in their intestinal cluster and found that they have higher expression in the anterior intestine cluster (new Figure S1C). These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9.

      The authors describe in the manuscript that they have performed INT1-specific RNAi for two C-type lectin genes that are upregulated during fasting. Due to a recent expansion of C-type lectin genes in C. elegans, there is a high chance of off-target effects of RNAi that is designed for members of this gene family. More trustworthy results could have been obtained using CRISPR-based loss-of-function alleles for these genes, one of which is publicly available. Also, the authors do not provide any explanation for why knockdown of these stress-response genes, which are activated in INT1 cells in response to food deprivation, results in improved resistance to pathogens. This, in fact, suggests a role of INT1 cells in increasing pathogen susceptibility, and not pathogen resistance, during food deprivation.

      We agree that RNAi targeting C-type lectin family members may be susceptible to off-target effects, and that validation with CRISPR null alleles, where available, would strengthen these findings. In the current study, we used INT1-specific RNAi as a cell-specific first-pass approach to test candidate gene function. We also agree that the pathogen phenotype requires cautious interpretation. Specifically, the finding that knockdown of fasting-induced INT1 lectin genes improves pathogen resistance does not support a simple protective model for these genes. Instead, it suggests that INT1-expressed stress-response genes modulate host susceptibility or host-pathogen interactions.

      Many of the studies in this field (e.g., references 2-4 in this article) have investigated the effects of food deprivation ranging from 4 hr to 24 hr, which results in activation of starvation responses in C. elegans. In contrast, the authors have used shorter time periods of fasting (30 min and 180 min), and most of their follow-up experiments have used 30 min of food deprivation. Previous work has shown that the effects of food deprivation can either accumulate over time (i.e., the effect gets stronger with longer food deprivation) or can be transient (i.e., only observed briefly after removal of food and not observed during long-term food deprivation). Starvation-induced transcription factors such as DAF-16/FoxO and HLH-30 show strong translocation to the nucleus only after 30 min of fasting. Though gene expression changes in all stages of food deprivation are of biological relevance, the authors have missed the opportunity to explore whether increased INS-7 secretion from the anterior intestine is dependent on these starvation-induced transcription factors (which can be easily tested using loss-of-function alleles) or is due to other fast-acting regulatory mechanisms induced due to the absence of food contents in the gut lumen. A previous study (PMID: 40991693) has shown that DAF-16 activation during prolonged starvation shuts down insulin peptide secretion from the intestinal epithelial cells. Hence, it is not clear if increased INS-7 secretion is only a feature of short-term food deprivation or is also a signature of long-term starvation (e.g., at 8 hr or 16 hr timepoints). Since most of the INS-7 secretion data in this study are for 30 min of fasting, it remains unknown whether the discovered regulators of INS-7 secretion can be generalized for extended food deprivation that triggers major metabolic changes, such as fat loss (e.g., conditions shown in Figure 1D).

      We agree that short-term food deprivation and prolonged starvation likely engage distinct regulatory mechanisms, and that our study primarily addresses an early phase of food deprivation rather than the full spectrum of starvation responses described in prior work. We selected the 30 min fasting condition because our previous study showed that INS-7 secretion is induced within this interval and returns to baseline upon refeeding, even before detectable intestinal fat loss. We also included a 180 min fasting condition to capture a later state associated with metabolic changes. However, we agree that the present study does not determine whether the regulators of INS-7 secretion identified here also govern secretion during more prolonged starvation (for example, 8 hr or 16 hr), nor does it test whether starvation-responsive transcription factors such as DAF-16 or HLH-30 contribute to this regulation. We appreciate that determining how this response transitions during prolonged starvation will be an important direction for future work.

      Two previous studies (PMID: 18025456, 40991693) have shown a strong reduction in the expression of ins-7 in the anterior intestine using GFP-based reporters (both promoter fusions and endogenous CRISPR-generated) and in whole-animal RNA-seq data from starved animals. These results are in contrast to the increased INS-7 secretion from INT1 cells during fasting that is reported in this study. The authors here have reported that INS-7 translation is higher in INT1 compared to INT2-9 during fed, acute fasted, and chronic fasted conditions, but they have not shown whether INS-7 translation is upregulated during acute and chronic fasting in INT1 cells in their TRAP-seq analysis. Knowing whether increased INS-7 secretion during acute fasting is due to increased transcription, translation, or secretion of INS-7 is crucial to resolve the discrepancy between these studies.

      In our dataset, INS-7 translation in INT1 tended to increase during acute fasting relative to the fed state (log<sub>2</sub>FC = 0.69), although this effect did not reach statistical significance after adjustment for multiple comparisons. Consistent with this trend, our secretion assay showed that INS-7 release from INT1 increases during fasting. However, we agree that the current data do not distinguish whether this increase in secretion is driven by enhanced synthesis, regulated release of pre-existing peptide stores, or a combination of both.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to understand whether the discrete segments of the C.elegans intestine were specialized to carry out distinct functions during an animal's exposure and adaptation to a fast-changing nutrient environment. To achieve this, the authors used a method called Translating ribosome affinity purification (TRAP), which provides a snapshot of what genes are being translated into proteins (and therefore functionally prioritized by the animal) under different fasting and re-feeding conditions. By expressing the TRAP constructs in two distinct segments of the intestine (INT1) and (INT2-9), the authors were able to identify how these segments responded to changing nutrient availability.

      Already under steady state nutrient conditions, the authors found that INT1 and INT2-9 appeared to have different 'tasks', with INT1 expressing more immune- and stress-response related genes. Exposing animals to different regimens of starvation and refeeding also showed marked differences between the intestinal segments, and the gene expression patterns in INT1 were consistent with INT1 cells playing an integrative role in linking nutrient cues to the secretion of insulin molecules that regulate fat metabolism with food intake. In summary, the data presented catalogue, for the first time, gene expression differences between two areas of the intestine, suspected to play different roles, and through clever experiments, links these gene expression changes to responses to nutrient availability.

      Strengths:

      The data presented catalogue - for the first time and in a careful manner - gene expression differences between two areas of the intestine. They strongly support the presence of intriguing differences between two areas of the intestine in immune, metabolic, and stress-response regulation, and link these gene expression changes to the responses of these regions to nutrient availability.

      Weaknesses:

      The conclusions of this paper are mostly well-supported by data, but the relevance of the changing gene expression patterns could be better clarified and extended in the discussion.

      We thank the reviewer for this constructive comment. In the revised manuscript, we have now expanded the discussion to more clearly interpret these dynamic translatomic changes in the context of intestinal subset specialization. The most pronounced difference between INT1 and INT2-9 cells is the enrichment of stress-response genes in INT1. Based on the present findings, together with our previous work identifying INS-7 as an INT1-secreted signal (PMID: 39127676), we propose that INT1 cells are sentinel enteroendocrine cells that integrate information from the luminal environment and the metabolic state of intestinal cells.

      Reviewer #3 (Public review):

      Summary:

      In this study, Liu and colleagues utilize TRAP-seq to profile the repertoire of actively translated mRNAs in different intestinal cell types (anterior INT1 vs. posterior INT2-9 cells) in C. elegans. A key goal of this study was to identify transcripts differentially expressed/translated between these intestinal cell subtypes in the context of animals being well fed or subjected to acute (30 minutes) or chronic (3 hours) starvation, followed by refeeding.

      The authors identify a number of differentially expressed genes across all of the conditions tested. They then provide an initial survey of the landscape of translatome changes through Weighted Gene Network Correlation Analysis (WGNA), and some high-level functional surveys via Gene Ontology (GO) term analysis and protein domain analysis. The authors validate the enriched expression patterns of some of their identified candidate genes using fluorescent promoter fusion reporters, confirming INT1-specific expression. The authors further implicate the role of several other candidate genes in pathogen avoidance and in response to nutritional cues by knocking them down specifically in INT1 cells by RNAi. Finally, the authors identify pyruvate as a major nutrient signal coming from the bacterial diet that suppresses the release of a key insulin peptide (INS-7), and identify some of the genes expressed in INT1 that are required for this response.

      Strengths:

      (1) Good use of and justification for TRAP-seq, because scRNA-seq would be difficult under the varied conditions used (starvation, refeeding).

      (2) The manuscript is generally clear to read, and the data are generally well-presented with good supporting data that includes replicates, sample sizes, error measurements, and associated statistics.

      (3) The dataset will be an interesting resource to mine for future studies focusing on mechanisms of how particular intestinal cell types respond to different environmental signals.

      Weaknesses:

      (1) A limitation of TRAP-seq, although powerful, is that only relative comparisons can be made between genotypes/conditions to identify differentially-expressed genes, rather than assessing whether a given gene is expressed at a certain level in a cell type under a certain condition. This limitation is due to the non-specific association of sticky RNA species with the beads during the immunoprecipitation step. This is a minor point, however, and the authors do a nice job of focusing their analysis on differentially expressed transcripts in the current study.

      We agree that a limitation of TRAP-seq is that it is best suited for relative comparisons across cell types or conditions, rather than for determining the absolute expression level of a given transcript in a specific cell type. As the reviewer notes, this limitation arises in part from nonspecific recovery of background or sticky RNAs during the immunoprecipitation step, complicating the interpretation of absolute expression levels. For this reason, our analysis was designed to focus primarily on differentially enriched transcripts between INT1 and INT2-9 cells and across feeding states, rather than on assigning absolute expression levels to individual genes. We appreciate the reviewer’s recognition of this point. Our study uses TRAP-seq specifically to define relative translatomic differences between intestinal subsets and physiological states, which is well aligned with the strengths of this approach.

      (2) Another limitation of the current study is that the experiments testing the role of candidate genes identified by their profiling experiments do not delve a bit deeper into providing a mechanistic understanding of the phenotypes being studied. At present, the results are thus viewed more as a genomics-based screen with some limited follow-up on interesting hits. However, this reviewer appreciates that when placed in the context of the work presented, a presentation of the profiling data along with some validation is an excellent starting point for future mechanistic studies elaborating on these interesting candidates.

      We agree that the current study does not fully resolve the molecular mechanisms by which the candidate genes identified by TRAP-seq regulate the phenotypes examined here. Our primary goal was to generate a spatially resolved translatomic framework for INT1 and INT2-9 cells across feeding states, and to perform focused validation of selected candidates to establish the physiological relevance of the profiling results. We therefore view the current functional analyses as an initial validation and proof of principle, rather than a comprehensive mechanistic dissection of the molecular pathways for each candidate. We appreciate the reviewer’s recognition that these findings provide an excellent starting point for future studies.

      Appraisal of whether the authors achieved their aims, and whether the results support their conclusions:

      The main goal of the study was to survey the dynamic responses at the level of actively translated mRNAs of the INT1 vs INT2-9 cells in response to metabolic challenge.

      Overall, the authors use established methods to perform their genome-wide analysis, and the set of differentially regulated genes is enriched for expected molecular functions and forms coherent networks in anticipated pathways.

      The validation experiments (promoter::GFP fusion reporters, INT1-specific knockdowns of highly regulated genes) further corroborate the quality of the TRAP-seq datasets generated.

      I have a few points for the authors that would further strengthen this work:

      (1) The authors rightfully focus on the top differentially-regulated candidates, but it's unclear at present how far down their fold change list would lead to expression pattern validations. It would be useful to test a few more promoter::GFP fusion reporters at different enrichment/fold-change/statistical cutoffs.

      Testing additional promoter::mNeonGreen reporters across a wider range of fold-change and statistical thresholds could be somewhat useful for calibrating ranked TRAP-seq candidate genes. However, given the variation in strains bearing extrachromosomal arrays, we did not consider this a stringent enough test, given that the sensitivity and dynamic range of RNA-seq far outpaces genetic fluorescence-based reporters. For these reasons, we focused on the top differentially enriched candidates to provide not only an initial validation of the dataset, but also to determine whether these candidates regulate biological functions in INT1 cells and thus serve as potentially useful biological readouts in future efforts.

      (2) Although the INT1-specific RNAi provides a convenient strategy for rapidly perturbing and testing genes of interest for phenotypes, independently validating the knockdowns with genetic mutants, or alternatively (if genes are essential), degron alleles.

      We agree that validating the INT1-specific RNAi phenotypes with independent genetic approaches, including null-allele or degron-based alleles for essential genes, would further strengthen the conclusions. In the current study, we used INT1-specific RNAi as a rapid and spatially restricted strategy to functionally test candidates identified by TRAP-seq and to determine whether these genes contribute to the specialized physiological functions of INT1 cells. We consider these experiments an initial validation of candidate function rather than a complete genetic dissection, which could be conducted in future efforts to study other aspects of INT1 function.

      Impact:

      The TRAP-seq data and list of differentially-expressed candidate genes will form an interesting set of high-priority candidates to study for their role in the reception and transduction of nutritional cues in response to food status and pathogens. This data will thus benefit the C. elegans community of researchers studying the mechanisms governing these phenomena.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments:

      (1) The authors need to describe the fasting method used in detail. Was fasting performed on unseeded NGM plates or in liquid (M9 buffer)? Were the animals washed with buffer prior to starvation? If yes, how many times? These details are critical for any researcher to follow up on their results.

      We have clarified the fasting/refeeding procedure in the revised Methods section. Briefly, worms were washed off OP50-seeded NGM plates with M9 buffer, washed three times in M9 buffer, and then transferred to unseeded NGM plates for fasting. For refeeding, worms were collected from the unseeded NGM plates with M9 buffer and transferred back to OP50-seeded NGM plates.

      (2) The authors claim that "INT1 and INT2-9 cells maintain fundamentally different molecular identities independent of any and all acute or chronic conditions", which they primarily based on Principal Component Analysis (PCA). The circles shown in Figure 2A are arbitrary, and many such circles can be drawn in the 2D space to separate the samples in different ways. The authors should show this comparison in a translatome-wide similarity heatmap with hierarchical clustering (similar to Fig.2C, but with all the experimental conditions and their replicates on both x- and y-axes).

      The ellipses shown in Figure 2A were generated using the stat_ellipse() function in ggplot2, which calculates the mean and covariance of the PC1 and PC2 for each line and draws ellipses corresponding to the 95% confidence level. The separation between lines is primarily driven by PC2, which accounts for 14% of the variance in the translatomic dataset. Although this difference is less pronounced when considering the full translatome, samples from the same line nevertheless cluster together, supporting line-specific differences in translatomic profile.

      (3) Figure 4 of the study shows 18 Venn diagrams for genes that are differentially expressed between INT1 and INT2-9 cell types in fed, fasted, and refed conditions. In the absence of any statistical comparisons, it is difficult to interpret whether the extents of overlap (higher or lower than expected) are significant. Ideally, P values for hypergeometric tests should be provided for the overlap regions.

      We appreciate the reviewer’s suggestion. We explored the use of hypergeometric testing, implemented through the SuperExactTest package in R, to assess the statistical significance of the overlaps shown in the Venn diagrams. However, this analysis yielded significant P values for essentially all overlap regions, including cases in which the degree of overlap was not especially informative and did not align with the interpretation presented in the text. This outcome likely reflects the dependence of the test on the size of the input gene sets and background universe, which can make statistical significance difficult to interpret meaningfully in this context.

      (4) The claims made in lines 242-244 (Figure 5D) need to be supported by P values from hypergeometric tests.

      Similar to the previous point.

      (5) The interpretation of Figures 6E and 6F described in lines 296-297 needs to be supported by statistical analyses. The authors claim that the undulating pattern of expression of the turquoise module genes is stronger in INT1 compared to INT2-9. However, based on Figures 2C and 6F, it appears that the expression change is not necessarily weaker in INT2-9, but instead is different, i.e., the expression of turquoise module genes goes up during fasting in INT1 and goes down after refeeding, while their expression goes up during fasting and stays up after refeeding in INT2-9 cells.

      We appreciate the reviewer’s point and agree that the turquoise module shows dynamic regulation in both cell populations. The key difference is not the presence versus absence of an undulating pattern, but rather the magnitude of that change, which is greater in INT1. Because of the limited number of biological replicates in some conditions, particularly the fasting group, we interpreted these results cautiously and used a nonparametric approach to assess differences in average module expression between states. This analysis indicated that the turquoise module changes significantly in both lines, but with a larger effect size in INT1. We have included the corresponding statistical analysis and effect size in the revised manuscript.

      (6) Since the INS-7 coelomocyte uptake assay was used extensively in this study, some representative microscopy images should be included to complement the quantification.

      We have added a new Figure 6G showing representative images corresponding to the quantification presented in Figure 6H.

      (7) The authors claim that INT1-specific fmo-2 RNAi results in reduced basal INS-7 secretion, but they do not have the direct statistical comparison for this. Were experiments shown in Figures 6G and 6K done on the same day?

      In the original Figure 6K (now Figure 6L), the data are presented as the percentage of normalized INS-7::mCherry fluorescence intensity relative to fed animals treated with vector RNAi. A statistical comparison between fed animals treated with INT1-specific fmo-2 RNAi and fed vector RNAi controls was performed and was significant. We have also clarified that the experiments shown in the original Figures 6G and 6I (now Figures 6H and 6J) were performed on the same day.

      (8) It is not clear why blocking the mitochondrial breakdown of pyruvate (Figures 7E and 7F) does not mimic the fasted state in terms of increased INS-7 secretion from INT1 cells. Doesn't this contradict the proposed model in this study? Can the authors speculate why this is the case?

      We do not interpret inhibition of pyruvate dehydrogenase or pyruvate carboxylase as equivalent to the fasted state. Rather, our model is that fasting induces INS-7 secretion by lowering intracellular pyruvate in INT1 cells. Under this framework, blocking mitochondrial pyruvate breakdown would be expected to reduce pyruvate utilization and thus maintain intracellular pyruvate, preventing the drop in pyruvate that normally occurs during fasting. This would explain why these manipulations suppress fasting-induced INS-7 secretion. To directly examine this possibility, we performed the experiment in Figure 7G, which tests whether maintaining pyruvate levels in INT1 cells during fasting is sufficient to suppress INS-7 secretion. The results are consistent with this interpretation and further support a model in which decreased intracellular pyruvate is a key determinant of fasting-induced INS-7 secretion.

      Minor comments:

      (1) Figures 2E and 2G are very similar and represent the same result in two different ways (unbiased vs guided comparison). One of these should be moved to the supplementary figures.

      Although these figures show similar patterns, they were derived from two independent analytical approaches, WGCNA and differential expression analysis. We therefore interpret the concordance between these independent methods as strengthening the robustness of the association and increasing confidence in the biological relevance of the observed pattern.

      (2) In Figures S1C, S1D, and S1E, a more significant P-value is shown with a smaller circle, and a less significant P-value is shown with a larger circle. This is confusing to the reader and should be inverted.

      We appreciate the reviewer’s comment and have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes used in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. Our further examination of marker expression across the intestinal clusters in the Ghaddar et al. (PMID: 37352352). dataset confirmed this limitation. These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9. We also note that the spatial identities of the intestinal clusters in Ghaddar et al. (PMID: 37352352) were not clearly established in the text or by spatial transcriptomic evidence, making it difficult to assign the annotated anterior, middle, and posterior clusters to specific intestinal cells. We have revised the manuscript accordingly and replaced the original figure panels.

      (3) Line 131 mentions the comprehensive characterization of the translatomic differences between INT1 and INT2-9 cells under each acute and chronic condition. However, the paragraph only discusses the differences in the fed condition. This is confusing, and the authors should mention the comparison between these cell types under acute and chronic conditions in subsequent sections where it is described.

      In this paragraph, we indeed discuss the ‘fed’ condition, but in subsequent sections we follow with details analyses of acute versus chronic, as well as regional differences across the intestine. We have clarified this in the opening sentence of the referenced paragraph.

      (4) The Venn diagrams in Figure 4 look very similar, and it is hard to differentiate between how 4A is different from 4I, how 4B is different from 4J, etc. The authors should include the labels for 'acute' or 'chronic' above each Venn diagram to guide the reader through these panels.

      We have added labels indicating the acute and chronic conditions to the left side of each Venn diagram in the revised Figure 4.

      (5) It is not clear in the figure legends how Figure 5E is different from Figure S6A, and how Figure 5F is different from Figure S7A. This should be better described in the figure legends.

      We have added a sentence to better describe this in the figure legend.

      (6) Figure 6C: Survival parameters such as median lifespan, number of animals for each condition, etc., should be reported for the different conditions.

      We have revised Figure 6C to indicate the number of animals analyzed in each condition, and the median survival for each group is now reported in the corresponding figure legend.

      (7) The colors used for control RNAi and clec-160 RNAi are very similar in Fig.6C. Easily distinguishable colors should be used.

      We have changed the colors as suggested.

      (8) The INT1-specific RNAi strain should be first described in line 285.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (9) Line 304: 'REF' should be replaced with the reference.

      We have replaced the “REF” with the reference (PMID: 39127676)

      (10) The P value for statistical comparison between the fed and 30 min refed states should be shown in Figures 6G, 6I, and 6K.

      We have now included the p value for the comparison as suggested. Figures 6G, 6I, and 6K are now labeled as 6H, 6J, and 6L, respectively.

      (11) In Figure 7, the authors should consider replacing the 'refed' label with 'recovery' because the pyruvate treatment was done in the absence of 'feeding' (= bacteria consumption).

      We appreciate the reviewer’s point. However, we chose to retain the label “refed” in Figure 7 to maintain consistency across the set of conditions examined, including 2% glucose and OP50 supernatant, which likewise do not involve bacterial consumption despite not showing effect on the refeeding response of INS-7 secretion.

      (12) The full form of DISN should be mentioned in the figure legend of Figure 7.

      We have included the full form of D1SN in the figure legend of Figure 7A.

      (13) Line 367: 'normalization' should be replaced with 'return to basal levels'. 'Normalization of INS-7 secretion' might also mean normalization of INS-7::mCherry signal to CLM::GFP signal.

      We have revised the wording per the reviewer's suggestion.

      (14) The methods section has a quantitative RT-PCR section, but it is not clear if RT-PCR data are reported in any of the figures. Also, no qPCR primers are listed in Table S3.

      We have removed the quantitative RT-PCR part from the methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors describe that the RPL-22-3xHA constructs are not integrated, at the very end, in the section "Limitations of the data". An earlier mention of this caveat would have been useful. In addition, it would help if the authors could provide their defense (which I think is very valid) of using non-integrated strains in the results section, as they describe the experimental setup. Also, some details were missing, which left me wanting to know: Were there expression differences? How were they accounted for? Was expression normalized between these two constructs, and if so, how?

      The RPL-22-3xHA construct was integrated into each line to ensure more consistent transgene expression across experiments. Because the INT2-9 construct is expressed in a larger number of cells than the INT1 construct, the INT2–9 samples yielded greater amounts of pulled-down nascent RNA, as reflected in the supplemental table and in the higher aligned RNA counts observed for the INT2-9 samples. To account for these differences, differential expression analysis was performed using DESeq2, which corrects for library size by estimating sample-specific size factors with the median-of-ratios method. Raw counts are then normalized using these size factors, thereby accounting for differences in sequencing depth and minimizing confounding effects due to variation in library size. Such differences are common in RNA-seq experiments, particularly when comparing samples derived from distinct input populations.

      (2) The 'acute' and 'chronic' exposures are thought through and carefully defined. The question I do have is whether the 3-hour fasting can be considered chronic fasting, given how surprisingly fast the animals lose their fat content. Could these kinetics indicate that the 30-minute fasting is reflective of mechanisms during which senses change in food availability, whereas the 30 minutes represents acute fasting (with chronic fasting - meaning fasting, during which the animal activated alternative pathways - occurring later)? While this may appear to be pure semantics, it could influence how the authors interpret their results. One method to more objectively separate an 'acute' from a 'chronic' stage may be to conduct a time course of fat loss-does fat loss plateau after 3 hours? The timing when the rate of decrease levels off could be more indicative of the beginning of a chronic phase.

      We appreciate this important point and agree that it should be more clearly discussed. We interpret acute fasting as a pre-fat-loss state, since it is 30 minutes off food and no difference in fat levels are detectable at this stage (Fig 1B). The translatomic changes observed under acute fasting therefore likely reflect food-sensing mechanisms and early preparatory responses that promote subsequent fat mobilization. In contrast, chronic fasting (180 minutes off food – see Fig 1D) appears to represent a post-fat-loss state, in which fat stores have already been depleted, and the corresponding translatomic changes likely reflect the effects of sustained metabolic stress.

      (3) The age of the animals used has to be more explicitly stated. Were these animals egg-laying? Or L4/young adults? This is likely to impact the changes that the animals undergo.

      Day 1 young adults were subjected to the fasting. Great care was taken to ensure consistency across biological replicates.

      (4) What is the rationale, in the authors' view, that stress response genes are apparently more enriched than metabolic or mitochondrial enzymes, and membrane receptor changes? Are the latter mostly regulated by PTMs/localization changes, etc?

      Based on our current data, we cannot exclude the possibility that metabolic or mitochondrial enzymes, as well as membrane receptors, are regulated in INT1 and INT2–9 cells through mechanisms not captured at the translatome level, including post-translational modification or changes in subcellular localization under different fasting and refeeding conditions.

      (5) The refeeding experiment with latex beads and killed OP50 is very clever. Details on when INS-7 was evaluated in the caoelomocytes would help the reader better understand and interpret these results.

      INS-7mCherry signal was evaluated in the coelomocytes immediately after refeeding; we included this information in the methods section and referenced our previous paper.

      (6) In the Discussion, I was looking for a more detailed context for how to think about the differences and similarities in the RNA-seq data between the two segments, and perhaps a discussion of whether there were any indications that the two segments communicated with each other.

      The data show that the most pronounced difference between INT1 and the rest of the intestine at the RNAseq level, is the expression of stress response genes in INT1. Although there are some nuanced differences, the prevalence of stress response genes persists across feeding and fasting conditions. This difference, combined with the evidence that INT1 cells secrete the enteroendocrine peptide INS-7 (Fig 6 and PMID: 39127676) is strongly reminiscent of the mammalian enteroendocrine cells, which also secrete peptides and show strong expression of stress response genes (PMID: 37626258 and 27148273). We suggest that this category term reflects not only a canonical stress response, but also a broader response to shifts in the luminal environment, which INT1 cells are anatomically poised to detect well before the absorption of nutrients has begun further down the intestine (INT2-9). Thus, we believe INT1 cells are a newly defined enteroendocrine cell type within the C. elegans intestine.

      Regarding communication between INT1 and INT2-9 – this is an intriguing possibility that we have considered, given that peptide genes and receptors are found in the RNAseq datasets. The extent to which the expression of these genes leads to functional effects is the subject of future investigation.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1A - It would be better to also show single fluorescent protein channels to assess the specificity of the expression patterns. A schematic or labels of where the INT1 vs. INT2-9 boundaries are located would be helpful to non-experts.

      (2) Figure 4 - At present, the Venn Diagrams are a very complicated way to visualize all of the comparisons/conditions. I would recommend that the authors consider using UpSet plots to better summarize the relevant comparisons they would like to make. The same consideration applies to Figure 5D.

      (3) Line 301 - Description of the INT1-specific RNAi strategy. I think it would be better to bring this information earlier, close to line 285, where the authors first mention performing INT1-specific RNAi experiments.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (4) Line 304 - I think the authors meant to cite a reference where the REF placeholder text is found

      We have replaced the “REF” with the reference.

    1. eLife Assessment

      This study investigated mitochondrial dysfunction and the impairment of the ciliary Sonic Hedgehog signaling in Lowe syndrome (LS), a timely topic given the limited research in this area. The data obtained from patient-derived iPSC neurons and a mouse model are solid. Although the main claims of the study are only partially supported by the current evidence, it provides a useful starting point for future functional studies investigating the link between mitochondrial defects and primary cilia in neural development.

    2. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how neural cell development is affected in Lowe syndrome. Using neural cultures differentiated from human iPSCs carrying either a LS mutation or a genetically engineered mutation in OCRL, the authors show a depletion of mitochondrial DNA and decrease in mitochondrial activities that correlate with an increased formation of astrocytes at the expense of neurons. Similar effects on mitochondria and on astrocyte development were observed in a LS mouse model. Moreover, these mutant brain cells are less likely to be ciliated and show a reduction in Sonic hedgehog signalling.

      Strengths/Weaknesses:

      The study derives strength from the analyses of two different models of Lowe syndrome, both reaching similar conclusions. However, the observed changes in mitochondrial defects, neuronal/astrocytic development and primary cilia are only correlated, with no attempt to investigate a causal relationship. Moreover, the mouse model is only analysed at the adult stage providing no insights into the development of the defects. Different brain regions are analysed with immunostainings and qRT-PCR making it challenging to draw clear correlations between these findings. The quality of the corresponding figures is often poor and the selection of markers is frequently inappropriate. Taken together, these limitations complicate the interpretations of the data and significantly limit the conclusions that can be drawn from the study.

      Although the study remains incomplete as main claims are only partially supported it can be used as a starting point for future functional studies into the link between mitochondrial defects and primary cilia in neural development.

      Comments on revised version:

      I am afraid the revised manuscript does little to address the concerns I raised in my initial review. The authors have primarily revised the text, removed over-interpretations and discussed critical points as limitations of the study. This gives the impression that key concerns have merely been rationalised, particularly as only a few new experiments are presented. My main concerns therefore remain:

      (1) The authors present three different phenotypes (altered neural differentiation, mitochondria dysfunction, alterations in primary cilia and ciliary Shh signalling) but a link between these phenotypes is not investigated. No mechanistic experiments are presented. Instead, the authors try to address the lack of a mechanism through refined wording but still use formulations that imply a direct link between these phenotypes. For example, their rebuttal letter finishes with the statement that the manuscript "provides a multi-model, cross-species framework linking mitochondrial dysfunction, ciliary signaling, and altered neural differentiation in Lowe syndrome". Similar formulations are used in the text.

      (2) The authors still claim that ciliary Shh signalling is reduced but ignore the fact that Shh mRNA in iN cells and Shh protein in the IOB mouse are significantly decreased. This reduction represents the most likely explanation for the reduced levels of Gli1 and Ptc1 mRNAs (Shh target genes), rather than dysfunction of cilia. In order to test for cilia dysfunction, the authors need to use experiments in which they quantify the response of control and OCRL mutant cells to exogenously added Shh protein or Shh agonists. Moreover, the increased Gli1 protein expression in the IOB mouse contradicts the reduced levels of Gli1 mRNA.

      (3) The analyses of the IOB mice are only done in 2 months old adult animals, nevertheless claims are made that changes in cell proportions are consequences of altered cell fate decisions. Alterations in proliferation and cell death are not addressed by experiments.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study investigated mitochondrial dysfunction and the impairment of the ciliary Sonic Hedgehog signaling in Lowe syndrome (LS), a timely topic given the limited research in this area. The data from patient iPSC-derived neurons and a mouse model were collected using solid methods, but the evidence supporting key claims is incomplete, and some technical aspects fall short of expectations. Despite these limitations, the study provides a useful foundation for exploring the relationship between mitochondrial defects and primary cilia in neural development.We appreciate the editorial assessment highlighting the importance of studying mitochondrial dysfunction and ciliary signaling in Lowe syndrome. We acknowledge that our study is largely associative, and we have revised the manuscript to clearly state this limitation, toned down causal claims, and emphasized that our work provides a foundation for future mechanistic studies.

      We appreciate the editorial assessment highlighting the importance of studying mitochondrial dysfunction and ciliary signaling in Lowe syndrome. We acknowledge that our study is largely associative, and we have revised the manuscript to clearly state this limitation, toned down causal claims, and emphasized that our work provides a foundation for future mechanistic studies.

      We have also:

      - Improved figure clarity and consistency

      - Corrected errors in gene annotations and normalization

      - Refined the mechanistic framework linking OCRL, mitochondria, and cilia

      New Experimental Data:

      Figure 5, Supplementary Figure 4. We confirmed mitochondrial defects by generating ocrl-KO zebrafish (Supplementary Figure 4). We first assessed mitochondrial reactive oxygen species (mitoROS) using MitoSOX staining. Next, we evaluated mitochondrial membrane potential (ΔΨm) using MitoTracker CMXRos. Finally, we assessed mitochondrial content via TOM20 staining. For all analyses, we focused on the ocular and cranial regions of the zebrafish to maintain consistency (see Author response image 1). Notably, previous studies have reported that ocrl-KO zebrafish exhibit seizures and brain developmental abnormalities, supporting their relevance as a model for Lowe syndrome-like phenotypes [1].

      Author response image 1.

      Figure 6 c. In addition to quantifying the proportion of ciliated cells in the IOB mouse brain, we measured cilia length and compared it between IOB and WT brain sections. Our results show that IOB mice exhibit elongated cilia compared to WT controls, suggesting that OCRL deficiency is associated with stress-related alterations in ciliary structure. These findings are consistent with previous studies reporting that cilia elongation can be associated with increased ROS levels and mitochondrial dysfunction [2,3].

      Public Reviews:

      Reviewer #1 (Public review):

      The preparation of the manuscript requires improvement. There are many errors in the presentation of data.

      We thank the reviewer for this important comment. We have carefully revised the manuscript to improve the clarity, accuracy, and consistency of data presentation.

      Specifically, we have corrected inconsistencies in gene nomenclature (e.g., CO2 vs COX2, DLOOP) across the text, figures, and legends. We standardized normalization methods and ensured consistency between figures and descriptions. We revised figure labels, legends, and annotations for clarity and accuracy. We corrected referencing errors and ensured appropriate citation of prior work. We improved overall figure quality and readability. In addition, we performed a thorough review of the entire manuscript to eliminate typographical errors and ensure consistency in terminology and data interpretation.

      The use of references needs to be re-considered. Sometimes a reference is used when in fact the results included in that paper are the opposite of what the authors intend.

      We thank the reviewer for this important comment. We have carefully re-evaluated all references throughout the manuscript to ensure that they accurately reflect the findings they are cited to support. In cases where the cited studies did not fully align with our interpretation or could be misleading, we have either revised the text to more accurately represent the original findings or replaced the references with more appropriate sources. We have also clarified instances where prior studies report differing or context-dependent results to avoid overinterpretation.

      The authors conclude the paper by claiming that mitochondrial dysfunction and impairments of the ciliary SHH contribute to abnormal neuronal differentiation in LS, but the mechanism by which this sequence of events might happen hasn't been shown.

      We thank the reviewer for this important comment. We agree that the current study does not establish a direct causal mechanism linking mitochondrial dysfunction, ciliary SHH signaling, and altered neuronal differentiation in Lowe syndrome. Our data demonstrate that these processes co-occur consistently across multiple model systems, supporting a potential functional relationship. However, we acknowledge that the precise sequence of events and mechanistic connections remains to be defined. To address this, we have revised the manuscript to clarify that our conclusions are based on associative findings rather than direct mechanistic evidence. We have also updated the Discussion to explicitly acknowledge this limitation and to frame our model (Figure 7) as a proposed working hypothesis. Future studies will be required to determine whether mitochondrial dysfunction directly impacts ciliary SHH signaling and how these pathways influence neuronal differentiation.

      Phenotype of increased astrocytes in both the IOB mouse brain or iPSC-derived cultures iN cells requires clarification as one of the markers used as an astrocyte marker, BRN2, is commonly used as a neuronal marker. As LS is a neurodevelopmental disorder, and the phenotype in question is related to differentiation, it is crucial to shed light on the developmental timeline in which this phenotype is seen in the mouse brain.

      We thank the reviewer for this important comment. We agree that the use of BRN2 as an astrocytic marker was inappropriate, as it is primarily recognized as a neuronal marker. Accordingly, we have revised the manuscript to remove BRN2 from the interpretation of astrocytic identity and now rely on GFAP expression as the primary astrocytic marker. We have also clarified this point in both the Results and figure legends to avoid misinterpretation. In addition, we have revised the text to more accurately describe our findings as an altered balance in neuronal versus astrocytic marker expression, rather than a definitive increase in astrocyte numbers.

      Regarding the developmental context, we acknowledge that Lowe syndrome is a neurodevelopmental disorder and that temporal aspects are highly relevant. In our study, the in vivo analyses were performed on adult 2-month-old IOB mouse brains, which we have now explicitly stated in the manuscript. We recognize that this limits our ability to directly assess developmental dynamics of lineage specification. We have therefore added this as a limitation in the Discussion and clarified that future studies examining earlier developmental stages will be necessary to determine when these alterations arise.

      Mitochondrial dysfunction in astrocytes has been shown to induce a ciliogenic program. However, almost the opposite is shown in this paper, with regards to ciliation. Morphology of the cilia was not assessed either, which is an important feature of ciliary homeostasis. The improper ciliary homeostasis here appears to be the improper Shh signalling, which has not been shown to be related to mitochondrial dysfunction. This leaves one wondering how exactly the different phenotypes shown in this paper are connected.

      We thank the reviewer for this important comment. We agree that the relationship between mitochondrial dysfunction, ciliogenesis, and Shh signaling is complex and not fully resolved in the current study.

      As noted by the reviewer, prior studies have reported that mitochondrial dysfunction can promote a ciliogenic program [4]. In contrast, our data show a reduced proportion of ciliated cells together with increased cilia length, indicating altered ciliary homeostasis rather than a straightforward increase in ciliogenesis. To address this point, we have revised the manuscript to describe our findings as context-dependent alterations in ciliary parameters more clearly, and we now explicitly discuss this apparent discrepancy with the literature in the Discussion. We also acknowledge the reviewer’s point regarding ciliary morphology. In the revised manuscript, we have included quantification of cilia length in addition to the proportion of ciliated cells, and we have expanded the Methods section to detail how these measurements were performed. We agree that additional ultrastructural and functional analyses would further strengthen the characterization of ciliary homeostasis, and we now include this as a limitation and future direction.

      Regarding the link between mitochondrial dysfunction, ciliary alterations, and Shh signaling, we agree that our study does not establish a direct mechanistic connection. Our data demonstrate that these phenotypes co-occur consistently across multiple models, but do not define causality. To address this concern, we have revised the manuscript to clarify that our conclusions are associative, and we now present our integrated model (Figure 7) as a working hypothesis rather than a demonstrated mechanism. We also explicitly state in the Discussion that future studies will be required to determine whether mitochondrial dysfunction directly impacts ciliary signaling and Shh pathway activity.

      This paper lacks a clear mechanistic approach. While the data validates the 3 broad phenotypes mentioned, there is a lack of connection between these phenotypes or an answer to why these phenotypes appear. While the discussion attempts to shed light on this by referencing previous studies, some of the referenced studies show contradicting results. Hence, it would be beneficial to clarify these gaps with further experiments and address the larger question of the connection between the mitochondria, Shh signalling, and astrocyte formation.

      We thank the reviewer for this important and insightful comment. We agree that the current study does not establish a direct mechanistic link connecting mitochondrial dysfunction, altered Shh signaling, and changes in neuronal versus astrocytic differentiation.

      Our primary goal in this work was to identify and validate phenotypes associated with OCRL deficiency across multiple independent model systems. We demonstrate that mitochondrial dysfunction, oxidative stress, altered ciliary/Shh signaling, and changes in neural lineage-associated markers co-occur consistently in these models. However, we acknowledge that the causal relationships between these processes remain to be defined.

      To address this concern, we have revised the manuscript to more clearly state that our conclusions are associative rather than mechanistic, and we now present our integrated model (Figure 7) as a working hypothesis that links these phenotypes through a potential mitochondria-ROS-signaling axis. We have also expanded the Discussion to explicitly acknowledge this limitation and to avoid overinterpretation of causality.

      In addition, we have carefully re-evaluated and revised the cited literature to ensure accuracy, particularly in cases where prior studies report context-dependent or seemingly contradictory effects of mitochondrial dysfunction on ciliogenesis and signaling pathways. These points are now discussed more explicitly to better position our findings within the existing literature.

      We agree that further experiments, such as targeted rescue of mitochondrial function or modulation of Shh signaling, will be necessary to establish causal relationships between these pathways. These directions are now clearly outlined in the revised Discussion as important next steps.

      Most importantly, there is no mention of how the loss of OCRL, a 5-phosphatase enzyme, results in the appearance of the mentioned phenotypes. Since there are multiple studies in the field of Lowe Syndrome that shed light on the various functions of OCRL, both catalytic and non-catalytic, it is important to address the role of OCRL in resulting in these phenotypes.

      We thank the reviewer for this important comment. We agree that the link between OCRL function and the observed phenotypes was not sufficiently developed in the original version of the manuscript.

      In the revised manuscript, we have expanded the Discussion to more clearly outline how loss of OCRL could contribute to the observed mitochondrial, ciliary, and differentiation phenotypes. OCRL encodes a PI(4,5)P₂ 5-phosphatase that regulates phosphoinositide homeostasis and membrane dynamics. Disruption of this activity is known to affect endolysosomal trafficking, actin organization, and membrane remodeling-processes that are critical for organelle maintenance and ciliary function. We now discuss how these alterations could impact mitochondrial homeostasis, for example, through defects in membrane contact sites, vesicular trafficking, or organelle quality control pathways.

      In addition, we have incorporated discussion of potential non-catalytic roles of OCRL, including protein–protein interactions and scaffolding functions, which may contribute to the coordination of intracellular trafficking and cytoskeletal organization. These aspects may provide an additional layer of regulation linking OCRL loss to both mitochondrial dysfunction and ciliary alterations.

      We emphasize that, while these mechanisms are supported by prior studies, our data do not directly test them. Therefore, we have carefully framed this section as a plausible mechanistic framework rather than a demonstrated pathway and have explicitly stated this limitation. We also outline future experiments aimed at dissecting catalytic versus non-catalytic contributions of OCRL to these phenotypes.

      There are numerous errors in the qPCR experiments performed concerning the genes that were assayed. The genes mentioned in the text section do not match those indicated in the graphs or legends. This takes away the confidence of the reader in this data.

      We thank the reviewer for this important observation. We agree that the inconsistencies between the genes described in the text and those shown in the figures and legends could reduce confidence in the data. In the revised manuscript, we have carefully rechecked all qPCR experiments and corrected the gene names across the Results, figures, and figure legends to ensure full consistency. We have also standardized the nomenclature throughout the manuscript (including consistent use of gene symbols and formatting) and verified that all plotted data correspond to the correct targets.

      In addition, we have clarified the qPCR methodology, including normalization (all data are normalized to GAPDH) and primer information, to improve transparency and reproducibility.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how neural cell development is affected in Lowe syndrome. Using neural cultures differentiated from human iPSCs carrying either an LS mutation or a genetically engineered mutation in OCRL, the authors show a depletion of mitochondrial DNA and a decrease in mitochondrial activities that correlate with an increased formation of astrocytes at the expense of neurons. Similar effects on mitochondria and on astrocyte development were observed in an LS mouse model. Moreover, these mutant brain cells are less likely to be ciliated and show a reduction in Sonic Hedgehog signalling.

      Strengths/Weaknesses:

      The study derives strength from the analyses of two different models of Lowe syndrome, both reaching similar conclusions. However, the observed changes in mitochondrial defects, neuronal/astrocytic development, and primary cilia are only correlated, with no attempt to investigate a causal relationship. Moreover, the mouse model is only analysed at the adult stage providing no insights into the development of the defects. Different brain regions are analysed with immunostainings and qRT-PCR making it challenging to draw clear correlations between these findings. The quality of the corresponding figures is often poor and the selection of markers is frequently inappropriate. Taken together, these limitations complicate the interpretations of the data and significantly limit the conclusions that can be drawn from the study.

      We have carefully revised the manuscript to address the concerns raised, and we have revised the manuscript with additional supporting data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors have checked the expression of neuronal markers NeuN and FoxG1, and apart from GFAP, they categorise Brn2 as one of the astrocytes markers that they have also checked. But Brn2 is not an astrocyte marker. It is a neuronal marker that is expressed in layer 2/3 of the cortex. In fact, Brn2 is reported to be a key driver of neurogenesis in primate telencephalon development1 and for reprogramming of astrocytes to neurons2. Hence, the only glial marker they have used here is GFAP. BRN2 is a neuronal marker. It has been used as a neuronal marker even in the reference (Zhang et al, 2013), from which the protocol for inducing iPSCs to induced neurons (iNs) was taken. Hence, the qPCR results in 1g of overexpression of BRN2 indicate an increase in expression of a neuronal marker, not an astrocyte marker.

      We thank the reviewer for this important and well-founded comment. We fully agree that BRN2 is a neuronal marker and not an astrocytic marker, and that its inclusion as an astrocyte marker in our original interpretation was incorrect.

      In the revised manuscript, we have removed BRN2 from the analysis and interpretation of astrocytic identity. We now treat BRN2 exclusively as a neuronal marker and have updated the Results, figure legends, and text accordingly. Specifically, the qPCR data previously presented in Figure 1g are now interpreted as reflecting neuronal marker expression, not astrocytic differentiation. We have also revised our conclusions to avoid overinterpretation of astrocyte abundance. Our findings are now described more accurately as an altered balance in neuronal versus astrocytic marker expression, rather than a definitive increase in astrocyte numbers. In this context, GFAP remains the primary astrocytic marker used in this study.

      We acknowledge the reviewer’s point that reliance on a single astrocytic marker is a limitation. This has now been explicitly stated in the Discussion, and we note that additional astrocyte markers will be required in future studies to more comprehensively define lineage-specific changes.

      Incorrect marker usage (BRN2 as astrocyte marker)

      We thank the reviewer for identifying this critical issue. We corrected the classification of BRN2 as a neuronal marker. Also, we re-analyzed the interpretation accordingly, revised all relevant text and figures. Importantly, Astrocyte conclusions are now based primarily on GFAP expression, and we explicitly acknowledge this limitation in the Discussion

      The graphs for the RT-PCR results indicate that gene expression values are normalized to actin whereas the legend mentions that they are normalized to GAPDH. This needs clarification.

      We thank the reviewer for pointing out this inconsistency. We confirm that all qPCR data were normalized to GAPDH, and the reference to actin was an error. This has now been corrected throughout the figures, legends, and text to ensure consistency.

      The use of wording to refer to the generation of induced neurons (iNs) should ideally be changed from "we developed"; as the protocol from Zhang et al, 2013 seems to have been directly adapted in this paper.

      We thank the reviewer for this helpful suggestion. We agree that the wording was inappropriate. In the revised manuscript, we have replaced “we developed” with language indicating that iNs were generated using an established protocol, and we now explicitly state that the method was adapted from Zhang et al., 2013 [5].

      OCRL KO iPSCs were obtained from Herbert Lachman's lab and not generated in this study. Hence, the Ran et al, 2013 reference is not necessary.

      We thank the reviewer for this clarification. We agree that the OCRL knockout iPSCs were obtained from Herbert Lachman’s laboratory and were not generated in this study. Accordingly, we have removed the Ran et al., 2013 reference and revised the manuscript to clearly state the origin of the OCRL KO iPSC line.

      The reference for Figure 1c is given as Ran et al, 2013 which is wrong. It should be Zhang et al, 2013.

      We thank the reviewer for noting this error. We have corrected the reference for Figure 1c from Ran et al., 2013 to Zhang et al., 2013 in the revised manuscript.Limited in vivo mitochondrial characterization

      Figure 1D, E: GFAP is a cytoskeletal marker but its expression here is very grainy and looks like an artifact. Is it possible to show the astrocyte phenotype using other astrocytes nuclei and cytosolic markers such as NFIA and S100B, respectively?

      We thank the reviewer for this important suggestion. We acknowledge that GFAP is a cytoskeletal marker and that the signal in the current images may appear granular. We have carefully re-evaluated the staining and image processing to ensure that the signal represents true GFAP expression and have improved the image quality and presentation in the revised figures. We agree that inclusion of additional astrocytic markers such as NFIA and S100B would further strengthen the characterization. While we were not able to include these additional markers in the current revision, we now explicitly acknowledge this as a limitation in the Discussion and note that future studies will incorporate a broader panel of astrocyte markers to more comprehensively define astrocytic identity.

      Since the authors have not used enough markers to understand the cell-type composition in WT and OCRL KO/mutant lines, it's not sufficient to conclude that the NPCs preferentially favour astrocytes over neuronal lineage. Any conclusive comments regarding the cell-state/cell-type specification necessitate evidence such as genetic lineage tracing using reporters for neuronal and astrocyte markers, and/or RNA/ATAC/scRNA sequencing.

      We thank the reviewer for this important point. We agree that the current marker panel is not sufficient to definitively determine cell-type composition or to conclude preferential lineage specification.

      In the revised manuscript, we have tempered our conclusions and now describe our findings as changes in neuronal versus astrocytic marker expression, rather than evidence of a shift in lineage fate. We also explicitly acknowledge this limitation in the Discussion. We agree that approaches such as genetic lineage tracing, reporter-based assays, and single-cell transcriptomic or epigenomic analyses (e.g., scRNA-seq or scATAC-seq) would be required to rigorously define cell-state transitions and lineage outcomes. These are important directions for future studies and are now highlighted in the revised manuscript.

      Figure 2:

      The mt-DNA gene CO2 was checked, not COX2. A typographical error in the written section, which does not match with the qPCR graph of the same.

      We thank the reviewer for noting this inconsistency. We confirm that the gene analyzed was CO2, and the reference to COX2 in the text was incorrect. This has now been corrected throughout the manuscript to ensure consistency between the text, figures, and qPCR data.

      The word "neurogenesis" is used very loosely throughout the paper. In the opinion of the reviewer, there is no evidence presented that there is a defect in neurogenesis in either of the models used in this paper.

      We thank the reviewer for this important comment. We agree that the term “neurogenesis” was used too broadly and is not directly supported by our data. In the revised manuscript, we have removed or replaced this term where appropriate and now refer more precisely to changes in neuronal versus astrocytic marker expression.

      Line 150: They say that they have examined the functional properties of mitochondria during neurogenesis but it would have been better to understand OXPHOS at various time points of neurogenesis to actually conclude reduced OXPHOS 'during neurogenesis'. Moreover, genes related to other pathways such as glycolysis could have been checked to understand the bioenergetics of LS patients. Also, oxidative stress could have been checked using more than one marker. Since they are trying to understand the functional role of mitochondrial defects during neurogenesis, they could have performed live imaging of mitochondrial potential during various stages of neurogenesis. Isolation of mitochondria from LS patients and transcriptomics/proteomics might provide further clues about mitochondrial defects.

      We thank the reviewer for these thoughtful suggestions. We agree that our data do not capture mitochondrial function across multiple stages of neurogenesis. In the revised manuscript, we have modified the wording to avoid implying temporal analysis “during neurogenesis” and instead describe mitochondrial parameters in differentiated cells. To strengthen the study, we have included additional in vivo validation in the zebrafish model, where we assessed multiple mitochondrial readouts, including mitochondrial membrane potential (ΔΨm) using MitoTracker CMXRos, oxidative mitochondrial stress ROS (mitoROS) using MitoSOX staining, and mitochondrial content (TOM20), supporting mitochondrial dysfunction across systems.

      Elevated astrocytic reaction during the differentiation of NSPCs in the Lowe syndrome (IOB) mouse model.

      Title: What does astrocyte reaction mean? This term should not be used without clear evidence of reactive astrocytes being present in the model.

      We thank the reviewer for this important comment. We agree that the term “astrocytic reaction” is not appropriate without specific evidence of reactive astrocytes. In the revised manuscript, we have removed this terminology and replaced it with more accurate wording, describing our findings as altered astrocytic marker expression. This change better reflects the data and avoids overinterpretation.

      In 1a, no quantification of the mouse brain size is given. From the given images alone, there appears to be no obvious decrease in brain size between the WT and IOB mice. This contradicts the text which indicates that the IOB mouse brain is smaller.

      We appreciate this important point. We have:

      Removed claims regarding reduced brain size

      Clarified that our analysis was limited to available sections and no definitive conclusion about global brain morphology can be made

      Figure 3E: Why is the astrocyte to neuron ratio measured using a cytoskeletal marker for astrocytes, GFAP but a nuclear marker for neurons, NeuN? Ratios to measure the percentage or proportion of astrocytes to neurons can only be checked by markers of the same nature such as GFAP to MAP2 (neuronal cytoskeletal marker) or NFIA (astrocyte nuclear marker) to NeuN.

      We thank the reviewer for this important point. We agree that comparing a cytoskeletal marker (GFAP) with a nuclear marker (NeuN) is not ideal for deriving cell-type ratios. In the revised manuscript, we have removed the astrocyte-to-neuron ratio analysis and now present these data as relative marker expression/signals rather than proportions. We have also clarified this limitation in the text and Discussion.

      (3) Lack of clarity in the experiments performed on the mice brains. PAX6 is used here as a neuronal marker along with NeuN, a mature neuronal marker. This is misleading as PAX6 is rather a marker for neural stem/progenitor cells and not neurons. The age of the mice has also not been mentioned, which is crucial considering the different markers used to characterize the mouse brain as well as since the authors are indicating that there is an abnormal neurodevelopment in the IOB mouse during development. Again, BRN2 is used here as an astrocyte marker. However, it is a neuronal marker. Hence, the phenotype of increased astrocytes currently is held by GFAP expression alone. Another astrocyte marker should be used.

      We thank the reviewer for these important points. We have revised the manuscript to correct marker interpretation, now describing PAX6 as a progenitor marker rather than neuronal, and BRN2 as a neuronal marker, removing it from astrocyte-related analysis. We have also explicitly stated the age of the mice (2 months) in the Methods and Results. In addition, we have tempered our conclusions, describing the data as changes in marker expression rather than definitive cell-type shifts, and we now acknowledge that reliance on GFAP as a single astrocytic marker is a limitation, which is discussed in the revised manuscript.

      Figure 6: Increase in astrocytes, mitochondrial dysfunction, and ciliary Shh signalling are 3 phenotypes discussed in this study. However, no experiments were done to shed light on the mechanistic connection between these phenotypes. This is reflected in the abstract shown in Figure 6. There is no comment on the mechanism behind these phenotypes.

      We thank the reviewer for this important comment. We agree that the current study does not establish a direct mechanistic link between mitochondrial dysfunction, altered ciliary Shh signaling, and changes in astrocytic markers. Our aim was to identify and validate these phenotypes across multiple models. In the revised manuscript, we have clarified that Figure 7 represents a proposed working model based on associative findings rather than a defined mechanism. We have also revised the Discussion to explicitly acknowledge this limitation and to outline future experiments required to establish causal relationships between these

      Reviewer #2 (Recommendations for the authors):

      The authors report interesting findings in two different experimental models but the manuscript would benefit significantly from an analysis of a potential causal relationship between different findings. They often mention neural stem cells or the neuron/glia switch but their analysis of the mouse mutant is restricted to the adult stage. A more consistent analysis of specific brain regions would also be beneficial.

      We thank the reviewer for this constructive comment. We agree that establishing causal relationships between the observed phenotypes is an important next step. In the revised manuscript, we have clarified that our conclusions are based on associative findings and have expanded the Discussion to outline experimental strategies that could address causality in future studies. We also acknowledge that our in vivo analysis is restricted to adult (2-month-old) IOB mouse brains, which limits our ability to assess developmental dynamics such as neural stem cell behavior or neuron-glial transitions. This limitation is now explicitly stated in the Discussion. Finally, we agree that region-specific analysis would strengthen the study. Due to the availability of samples, our analysis was not systematically performed across defined brain regions. We now acknowledge this limitation and note that future studies focusing on specific regions (e.g., cortex, hippocampus) will be important to better understand the spatial aspects of the phenotype.

      Figure 1: The authors only measured the expression of marker genes by qRT-PCR. This could reflect higher expression levels in individual cells rather than a change in the proportion of neurons and astrocytes. They need to determine the cell proportions of astrocytes and neurons in addition. Moreover, there is a poor marker choice. Loss of FOXG1 expression could indicate a loss of telencephalic identity. BRN2 is expressed by cortical neurons.

      We thank the reviewer for this important comment. We agree that qPCR-based marker analysis does not directly reflect cell-type proportions and may instead represent changes in gene expression per cell. Accordingly, we have revised the manuscript to avoid conclusions about cell proportions and now describe the data as changes in marker expression. We have also corrected marker interpretation, removing BRN2 from astrocyte analysis and clarifying that FOXG1 reflects telencephalic identity. These limitations and the need for more comprehensive cell-type characterization are now acknowledged in the Discussion.

      In Figure 2, the authors determine the properties of mitochondria and claim that functional mitochondrial activities are decreased during neurogenesis in mutant iN cells. They need to take into account that according to Figure 1 the proportion of neurons and astrocytes may be changed. Hence, the decreased mitochondrial activity may reflect a fundamental difference between neurons and astrocytes. The authors need to clearly distinguish between neurons and astrocytes in their analysis. Moreover, the use of the term neurogenesis is confusing. They are analysing the neuron-to-glial switch, not the formation of neurons.

      We thank the reviewer for this important comment. We agree that differences in cell-type composition may influence mitochondrial measurements. In the revised manuscript, we have tempered our interpretation, describing these data as changes in mitochondrial parameters at the population level rather than neuron-specific effects. We also acknowledge this limitation in the Discussion and note that cell-type-specific analyses will be required in future studies. In addition, we have revised the terminology throughout the manuscript, removing the term “neurogenesis” and instead referring to changes in neuronal versus glial marker expression to more accurately reflect the scope of our analysis.

      Figure 3: The authors claim that astrocyte numbers are elevated in the IOB mouse model, however, it seems as if the authors analysed late postnatal, potentially adult brains but no age of the brains is provided. Given the large time lag between the formation of astrocytes and their analysis, the increased number of astrocytes could be due to a number of processes including altered proliferation and cell death. The authors need to investigate the proportion of astrocytes and neurons closer to the neuron-to-glia switch. Cell fate experiments like the long-term application of BrdU would be much better suited and would provide mechanistic insights. Again, markers are not adequate to reach their conclusion. Pax6 is only expressed in a tiny subset of neurons, Brn2 on the other hand is not astrocyte-specific as it is expressed in cortical neurons as well. Moreover, qRT-PCR analyses were done in the cortex and hippocampus whereas the boxes in Figure 3D are located in the basal ganglia. It would be much more informative and provide better comparisons to perform gene expression analysis and cell counts in the same brain regions.

      We thank the reviewer for these important and constructive comments. We agree that our analysis is limited by the use of adult (2-month-old) IOB mouse brains, which do not allow direct assessment of developmental processes such as the neuron-to-glia transition. We have now explicitly stated the age of the animals and clarified this limitation in the Discussion, including the possibility that changes in astrocytic markers may reflect processes such as proliferation or survival rather than lineage specification.

      We also agree that our marker panel was insufficient for definitive conclusions. Accordingly, we have revised the manuscript to remove overinterpretation, corrected marker usage, and now describe the data as changes in marker expression rather than cell-type proportions. The need for more rigorous approaches, such as lineage tracing (e.g., BrdU) and expanded marker panels, is now acknowledged as a future direction. Finally, we thank the reviewer for pointing out the inconsistency in the brain regions analyzed. We have clarified the regions used for qPCR, and we now explicitly acknowledge this limitation, noting that future studies will aim to perform region-matched molecular and histological analyses for more accurate comparisons.

      Experiments in Figure 4 assess "whether changes in mitochondrial activity are involved in the altered differentiation of stem cells and progenitor cells in the LS mouse model" in 3-month-old brain sections. The murine adult brain only contains a few neural stem cells in the SVZ and in the dentate gyrus. Instead, this analysis needed to be done at late embryonic/early postnatal stages to capture the neuronal/glial switch. In addition, RT-PCR and immunostainings should be performed in the same brain region as stated above.

      We thank the reviewer for this important point. We agree that analysis in adult (2-month-old) brains does not capture developmental stages such as the neuron-glia transition. We have revised the manuscript to remove implications of developmental analysis and now describe these data as mitochondrial parameters in adult tissue, explicitly acknowledging this limitation in the Discussion. We also clarify the brain regions used for qPCR and immunostaining and note as a limitation that these were not fully matched; future studies will perform region-specific, developmentally timed analyses.

      Figure 5: The authors examine a potential link between mitochondrial defects and primary cilia. Mutant iN cell cultures contain lower levels of SHH mRNA and show concomitantly lower expression of the SHH target genes GLI1 and PTCH1. The authors link this finding with a reduced proportion of ciliated cells but the reduced SHH signalling is most likely explained by the decreased SHH expression. The authors also limit their analysis of primary cilia to one brain region, but they should also include the cortex and hippocampus as these regions were used for their qRT-PCR analysis. SHH signalling acts as a switch to stop the proteolytic processing of GLI3 and to promote the formation of the GLI3 activator form. It is therefore important to determine the ratio of GLI3 repressor and GLI3 activator using western blots. The authors claim that they found defective cilia formation, but cilia are poorly characterised. Are there differences in intraflagellar transport, the formation of the transition zone, etc? Is ciliary length altered? The authors only make a correlative link between mitochondrial defects and cilia but present no experiments to investigate causation. They should at least discuss potential mechanisms which could explain defects in cilia.

      We thank the reviewer for these insightful comments. We agree that reduced SHH pathway activity may be influenced by decreased SHH expression, and we have revised the text to avoid overattributing this effect to ciliary changes. Our conclusions are now framed as associative, not causal.

      We have expanded our cilia analysis to include quantification of both the proportion of ciliated cells and cilia length, and clarified these methods in the manuscript. We also acknowledge that additional characterization (e.g., intraflagellar transport, transition zone structure, GLI3 activator/repressor ratios) would further strengthen the analysis, and we now include this as a limitation and future direction. Regarding regional analysis, we agree that broader brain region coverage would be valuable. Due to sample availability, our analysis was limited, and this is now explicitly acknowledged as a limitation, with future studies aimed at region-matched analyses (e.g., cortex and hippocampus). Finally, we have expanded the Discussion to outline potential mechanisms linking mitochondrial dysfunction and ciliary alterations, while clearly stating that causal relationships remain to be established.

      (1) The methods section does not contain any information on how immunostainings on brain sections were performed.

      We agree that the description of immunostaining on brain sections was missing. We have now added a detailed protocol for brain section immunostaining in the Methods section to improve clarity and reproducibility.

      (2) The abbreviation "RT-PCR" is used for both, real-time PCR and reverse transcription PCR

      We also acknowledge the inconsistent use of the term “RT-PCR.” In the revised manuscript, we have standardized the terminology, using “qPCR” (quantitative real-time PCR) throughout to avoid confusion

      Conclusion

      We believe that these revisions significantly strengthen the manuscript. While the study remains primarily associative, it provides a multi-model, cross-species framework linking mitochondrial dysfunction, ciliary signaling, and altered neural differentiation in Lowe syndrome.

      References:

      (1) Ramirez IB-R, Pietka G, Jones DR, Divecha N, Alia A, Baraban SC, et al. Impaired neural development in a zebrafish model for Lowe syndrome. Hum Mol Genet. 2012;21:1744–59. https://doi.org/10.1093/hmg/ddr608

      (2) Kim JI, Kim J, Jang H-S, Noh MR, Lipschutz JH, Park KM. Reduction of oxidative stress during recovery accelerates normalization of primary cilia length that is altered after ischemic injury in murine kidneys. Am J Physiol Renal Physiol. 2013;304:F1283-1294. https://doi.org/10.1152/ajprenal.00427.2012

      (3) Moruzzi N, Valladolid-Acebes I, Kannabiran SA, Bulgaro S, Burtscher I, Leibiger B, et al. Mitochondrial impairment and intracellular reactive oxygen species alter primary cilia morphology. Life Sci Alliance. 2022;5:e202201505. https://doi.org/10.26508/lsa.202201505

      (4) Ignatenko O, Malinen S, Rybas S, Vihinen H, Nikkanen J, Kononov A, et al. Mitochondrial dysfunction compromises ciliary homeostasis in astrocytes. J Cell Biol. 2022;222:e202203019. https://doi.org/10.1083/jcb.202203019

      (5) Zhang Y, Pak C, Han Y, Ahlenius H, Zhang Z, Chanda S, et al. Rapid Single-Step Induction of Functional Neurons from Human Pluripotent Stem Cells. Neuron. 2013;78:785–98. https://doi.org/10.1016/j.neuron.2013.05.029

    1. eLife Assessment

      This study introduces an artificial-intelligence tool that estimates fat in skull bone marrow from routine brain scans, enabling large studies that were previously impractical. The evidence for the method's repeatability and for identifying genetic links is convincing overall. The authors identify genes, diseases, and other biological characteristics that are linked to skull marrow fat, which represents an important advance. The work will be of most interest to researchers using large imaging biobanks and those studying ageing-related changes across bone, blood, and brain.

    2. Reviewer #1 (Public review):

      The authors of this study developed a method to quantify calvarial bone marrow from MRI head scans, enabling study of its composition in large datasets of adults, usually collected to study the brain. Bone marrow intensity can be semi-quantitatively measured in T1-weighted MRI scans due to the greater signal intensity of fat than watery red marrow. This is an ingenious use of the MRI-produced information for other important phenotypes, such as bone structure and marrow content. Different head types were tested for complying to the model, which is notable.

      The model was also successfully validated using several publicly available MRI resources - real data - in (1) dataset consisting of 30 individuals that were scanned 10 times each at 3-day intervals, and (2) the monozygotic (MZ) twin data from the Human Connectome Project cohort. Then the authors applied this validated method to head-MRI scans from the UK Biobank (n=33,042) to extract information on spatial distribution of bone marrow adiposity (BMA) in the calvaria, allowing a GWAS to identify associated genes.

      The authors revealed high heritability and identified 41 genetic loci significantly associated with the BMA trait, including six sex-specific loci. Of note, statistics estimate that 99% of BMA trait-influencing variants are shared with BMD (497 of 500 variants), which may mean these results demonstrate the biological relevance to bone health. Some of the BMA genes were found related to the Wnt pathway, including WNT16, WNT4, NXN; this is a "positive control", since the Wnt/β-catenin signaling pathway was suggested as an important determinant of BMA. Also, associations in genes (BMP4, DLX5, LGR4, LRP4, SFRP4) that are known to specifically influence adiposity, are encouraging. Integrating mapped genes with bone marrow single-cell RNA-seq data revealed patterns of adipogenic lineage differentiation and lipid loading.

      The study also investigated genetic overlap between BMA and twelve (or 13) "brain and body" traits, and identified significant genetic correlations with BMI, cognitive ability and Parkinson's disease.

      In sum, since MRI head scans present a hitherto unexplored opportunity to address unresolved aspects of bone marrow biology, this study is both timely and innovative.

      Comments on revised version:

      The authors responded most of this reviewer's comments. Their explanations are convincing. Yet, upon re-reading the revised version of this paper, I still have concerns about the clarity of mostly analysis presentation, e.g.:

      Line 130-133: the sentence is still unclear: "To obtain the BM signal intensity for an individual datapoint of the calvarium, we ... averaged these BM intensities to get the (average?) BM intensity for that datapoint. Then, we averaged (again?) these datapoint intensities across the calvarium to produce the global BMA measure for the scan."

      Also, I still cannot understand whether the "overlap between the true and predicted bone marrow ...below 0.7" is concerning or not, - whether this threshold of 0.7 is arbitrary.

      Genetic correlation: pls. make sure it's clear that the Rg was calculated using SNP "effect sizes".

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to enable large-scale measurement of fat in skull bone marrow using routine structural brain MRI scans. They present a neural-network pipeline trained largely on realistic simulated examples and show that the resulting skull marrow measure is highly repeatable in test-retest data and consistent in monozygotic twins. Applying it to ~33,000 UK Biobank participants, they report expected population patterns (including sex- and menopause-related differences) and identify genetic and health-related associations, creating an important resource that can be built upon by researchers interested in BMA, imaging, bone, metabolism, neuroscience, ageing, haematology, and other fields.

      Major strengths:

      A notable methodological strength is the training strategy: by using a large, simulated dataset that captures plausible variation in skull-layer thickness and MRI intensity, the authors reduce reliance on scarce expert-labelled images. The modelling choice (using 1D intensity profiles through the skull rather than analysing the full 3D volume) appears well matched to the anatomy and offers an efficient approach for thin, layered structures. Multiple validation steps (including test-retest reliability and twin concordance) support the robustness of the measurement pipeline.

      On the biological and genetic side, the study demonstrates that the skull BMA estimate relates to known correlates of marrow fat (e.g., age/sex/menopause patterns/bone density) and integrates population imaging with large-scale genetic analysis to highlight loci and candidate genes with plausible relevance to skeletal and marrow biology. The inclusion of cross-ancestry analyses and integration with cell-type-resolved gene-expression resources further improves interpretability and usability for the community.

      Major limitations:

      The main limitation is conceptual rather than technical: the phenotype is derived from T1-weighted MRI intensity, which does not directly separate fat and water signals and can vary with scanner and sequence settings. The manuscript provides convincing evidence that the measure is reproducible and biologically meaningful, but it should still be interpreted as a semi-quantitative proxy for marrow fat rather than a direct fat-fraction measurement. Accordingly, the genetic and phenotypic associations are likely informative, but the most direct claims about "adiposity" would be stronger if anchored to established quantitative fat-measurement imaging or spectroscopy in the skull.

      The genetic "replication" analysis in a smaller, ancestrally heterogeneous non-European-ancestry sample is useful as a test of transferability, but it is not equivalent to replication in an independent cohort of similar ancestry and is expected to show reduced SNP-level reproducibility because of differences in sample size and genetic background. This should be clearly framed so readers understand what level of generalisation is supported by the current evidence.

      Likely impact and utility:

      Overall, the work provides a practical method for extracting new biological information from widely available brain MRI scans and should be particularly useful to researchers working with large imaging biobanks and those studying connections between bone, blood, metabolism, and brain ageing. The combination of a scalable measurement approach and openly reported genetic results is likely to accelerate follow-up studies, including cross-cohort comparisons and mechanistic work on candidate pathways.

    4. Reviewer #3 (Public review):

      Summary:

      This paper addresses a fundamental gap in bone biology: our near-complete ignorance of the in vivo dynamics of calvarial bone marrow adiposity (BMA) at population scale. The authors developed an elegant artificial neural network trained on simulated data to automatically localize and quantify the bone marrow layer within standard T1-weighted MRI head scans; scans originally acquired to study the brain but harboring rich, unexploited information about adjacent bone. Applying this method to over 33,000 individuals from the UK Biobank, they accomplished three things that had never been done before: (1) they precisely quantified the sex-dimorphic age trajectory of calvarial BMA, including the dramatic post-menopausal rise and the protective role of hormone replacement therapy; (2) they performed the first well-powered GWAS of this trait, identifying 41 genome-wide significant loci including six sex-specific ones, with SNP heritability of 31.5%; and (3) they revealed significant genetic correlations and overlap between BMA and traits including bone mineral density, Parkinson's disease, and general cognitive ability, a finding made all the more intriguing by the recently described direct vascular channels connecting calvarial bone marrow to the meninges. Integration of GWAS genes with single-cell RNA-sequencing data from mesenchymal lineage cells further illuminated which genes govern lineage commitment to the adipogenic pathway versus lipid loading in mature adipocytes.

      Comments on revised version.

      The reviews raised substantive points across three domains, and the authors engaged with every one of them seriously and thoroughly.

      On the validation of T1-weighted MRI as a measure of BMA: Reviewer 2 raised the strongest concern, arguing that T1-weighted signal intensity had never been formally validated as a quantitative fat-fraction measure in the calvarium. The authors responded with both a principled scientific argument and new data. They assembled existing literature demonstrating that T1-weighted signal is an established semi-quantitative proxy for marrow fat in multiple skeletal sites (Loevner et al. 2002, Shen et al. 2013, Zhang et al. 2020), and provided additional comparative analyses against quantitative T1 relaxation maps, multiple intensity normalization strategies (KDE, WhiteStripe, GMM, FCM, Z-score), DEXA-derived bone mineral density, and osteoporosis status. The biological coherence of their findings, recapitulating known sex and age profiles, identifying genes already established in cell and animal models of BMA biology, and estimating heritabilities consistent with twin data constitutes powerful, convergent evidence for construct validity. Their point that a semi-quantitative measure of a highly variable, well-demarcated biological signal can outperform a perfectly precise measure of a poorly defined entity is methodologically sound and well-argued.

      On sex differences and the role of Hyperostosis frontalis interna: Reviewer 1 raised the clinically astute concern that Hyperostosis frontalis interna (HFI), a condition of inner table thickening prevalent in up to 49% of postmenopausal women, could confound calvarial BMA measurements and drive apparent sex differences. The authors performed a dedicated new analysis, stratifying BMA-BMD associations by sex and age group. They demonstrated that (1) the BMA-BMD association is robust in both males and females, (2) it remains stable across age groups, and (3) the neural network trained on simulations incorporating wide anatomical variation including inner table thickness is inherently resistant to moderate inner table thickening. Given that HFI is restricted to the frontal bone, which represents only a fraction of the calvarial surface, and that severe cases are rare (ICD-10 prevalence ~0.02% in the UK Biobank), the authors make a convincing case that this does not materially bias their results. Their suggestion that the method could itself be used in future work to study the genetic architecture of HFI is a nice forward-looking addition.

      On genetic correlation interpretation and cross-trait pleiotropy: Reviewer 1 asked for clarification of the vertical versus horizontal pleiotropy distinction and for formal Mendelian randomization to support the possible causal effect of BMA on cognition. The authors appropriately clarified the conceptual framework in the revised text and, rather than overstating a causal claim without the supporting analysis, responsibly softened the language to "may be consistent with the hypothesis that BMA could have a causal effect on cognition." This is scientifically honest and appropriate.

      On mouse scRNAseq and its relevance to humans: The authors acknowledged that the results section had not explicitly stated the mouse origin of the scRNAseq data, corrected this, and provided a well-justified rationale for the relevance of mouse mesenchymal lineage data to human BMA biology, which is a well-established and widely accepted model system in this field.

      On GWAS replication: The claim that the study lacked replication was addressed by clarifying the a priori separation of discovery (white British, n=33,042) and replication (non-white British, n=4,958) samples, with 62% of significant discovery SNPs and 95% of lead SNPs replicating in the correct direction.

      Overall Assessment:

      This is a technically innovative, scientifically rigorous, and biologically meaningful paper. The method is genuinely novel, the sample size is among the largest ever applied to this phenotype, the genetic findings are well-powered and well-replicated, and the integration across imaging, genetics, and single-cell transcriptomics is exemplary. The authors have engaged with every substantive reviewer criticism in good faith, producing new analyses where appropriate and defending, and convincingly, with findings that were challenged without adequate basis. The revised manuscript is strengthened throughout.

      This paper opens a new window quite literally, through the skull - into bone marrow biology at a scale and resolution that has never been achieved before.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors of this study developed a method to quantify calvarial bone marrow from MRI head scans, enabling the study of its composition in large datasets of adults, usually collected to study the brain. Bone marrow intensity can be semi-quantitatively measured in T1-weighted MRI scans due to the greater signal intensity of fat than watery red marrow. This is an ingenious use of the MRI-produced information for other important phenotypes, such as bone structure and marrow content. Different head types were tested for complying with the model, which is notable.

      The model was also successfully validated using several publicly available MRI resources - real data - in (1) a dataset consisting of 30 individuals that were scanned 10 times each at 3-day intervals, and (2) the monozygotic (MZ) twin data from the Human Connectome Project cohort. Then the authors applied this validated method to head-MRI scans from the UK Biobank (n=33,042) to extract information on the spatial distribution of bone marrow adiposity (BMA) in the calvaria, allowing a GWAS to identify associated genes.

      The authors revealed high heritability and identified 41 genetic loci significantly associated with the BMA trait, including six sex-specific loci. Of note, statistics estimate that 99% of BMA trait-influencing variants are shared with BMD (497 of 500 variants), which may mean these results demonstrate the biological relevance to bone health. Some of the BMA genes were found related to the Wnt pathway, including WNT16, WNT4, NXN; this is a "positive control", since the Wnt/β-catenin signaling pathway was suggested as an important determinant of BMA. Also, associations in genes (BMP4, DLX5, LGR4, LRP4, SFRP4) that are known to specifically influence adiposity, are encouraging. Integrating mapped genes with bone marrow single-cell RNA-seq data revealed patterns of adipogenic lineage differentiation and lipid loading.

      With regards to the reviewer’s comment on the overlap between BMA and BMD trait-influencing variants, we would like to add that the correlation of effect sizes within the overlap is -0.95 as we would expect from bifurcating differentiation of mesenchymal stem cells into osteoblasts or adipocytes: the underlying common biology driving these traits results in shared traits with a negative correlation in the effects.

      The study also investigated the genetic overlap between BMA and twelve (or 13) "brain and body" traits and identified significant genetic correlations with BMI, cognitive ability, and Parkinson's disease.

      In sum, since MRI head scans present a hitherto unexplored opportunity to address unresolved aspects of bone marrow biology, this study is both timely and innovative.

      There are, however, some assumptions, findings, and their interpretation, which require more critical focus.

      Sex-specificity is well described and studied here. Men have higher BMA than women, but post-menopausal women catch up in the BMA values. The authors believe that calvarial marrow has a number of features that make it particularly well-suited to the study of BMA process - which is clinically important in other bone sites. It has a simple "sandwiched" structure that they are able to model. This is true only to some extent: a condition called "Hyperostosis frontalis interna", of unknown etiology (described by Smith & Hemphill in 1956) - is characterized by irregular overgrowth of the inner table of the frontal bone (symmetric/bilateral). Although not of clinical significance, typically benign, studies report a prevalence of 12%; However, it's most common in postmenopausal women - where prevalences up to 49% in women over the age of 65 - have been reported. Thus, sexual dimorphism is obvious and the effect of estrogen is likely shared with whichever bone - and marrow - age-related pathology. So, for women not using HRT, this new layer of the bone might interfere with the calvarial BMA readings and in turn, affect the BMA-related analyses.

      Thank you for bringing the "Hyperostosis frontalis interna" condition to our attention. It is particularly interesting to hear that the etiology is unknown and one may suspect that some kind of calvarial bone marrow dysregulation may be part of the cause. Our model for bone marrow location was trained on simulated data which included variation in the thickness of all anatomical layers (including the inner table), so it will be robust to some thickening of the inner table. It might not be robust to the most extreme cases of inner table thickening (as described in some case reports), but these are rare. Further, it should also be noted that the other calvarial bones, representing a much greater fraction of the calvarial surface, remain largely unaffected by the thickening and would therefore yield correct localisation of the bone marrow layers. In summary, although the severe cases of hyperostosis frontalis interna have the potential to affect our identification of the bone marrow layer, the low frequency of such cases and the restriction of the phenotype to the frontal bone means that the potential for bias is very limited.

      It would be interesting to develop a method for detection of thickened inner bone so that the condition’s prevalence can be quantified in a large sample like the UK Biobank and its genetic architecture be determined. This could help elucidate the etiology.

      The authors suspect that the effect of BMA on BMD may be biased in women; they should comment on those "with low BMD and high BMA" given that hyperostosis frontalis might be an issue. A strong effect of SNPs in the ESR1 chromosomal region might be akin to the above concern.

      Thank you for raising this point, which we have followed up with a new analysis.

      According to ICD-10 data in UK Biobank there are only N=105 individuals with an M85.2-diagnosed disorder. Given the total sample size of N=446,814 individuals with ICD-10 data, this would translate to a prevalence of 0.02%, which speaks for an underdiagnosis in this sample such that we cannot simply remove diagnosed individuals to control for a potential diagnostic confound.

      We have therefore taken a different approach to investigate this potential issue: As you elaborated in your previous comment, the prevalence of hyperostosis frontalis increases with age in females. The literature also suggests that prevalence rates do not differ between males and females in young age / prior to menopause. Therefore, we have studied the association between BMD and BMA for males and females separately, and in two age groups based on a median split of our sample (left plot: younger than 65, right plot: subjects older than 65). In these plots, the relatively large shift in female BMA and BMD is visible with the large yellow cloud at low BMA and high BMD in the left plot disappearing in the right plot. Despite this, we observe:

      (1) Associations in both males (blue) and females (yellow), suggesting that the associations were not driven only by females.

      (2) BMA-BMD association is largely similar across the two age groups.

      If we consider that the old age group is likely to contain more cases of hyperostosis frontalis than the young group, and if we consider that old-aged females are more likely to be in this condition than men of any age, then we would expect an impact of hyperostosis frontalis on our measures to result in observable differences in Author response image 1. This is not the case. We see global age-related shifts in BMA in women, yet the association with BMD remains similar across age groups.

      The technical properties of our neural network (trained on simulated data) makes it unlikely that frontal bone will contaminate the bone marrow detection globally (description above) and these results show that hyperostosis frontalis is not a considerable issue in our analysis.

      Author response image 1.

      Then, there is a perfect overlap of the BMA SNPs that are shared with BMD (497 of 500 variants), which may prove a "face validity" of the MRI-derived BMA. However, the BMD in the study was heel-derived eBMD - which is a good proxy for osteoporosis and is mostly driven by trabecular bone. Thus, there might be a concern that the BMA metrics capture some trabecular BMD.

      The reviewer is correct in pointing out that the BMA causal variants are a near-perfect subset of the BMD causal variants. The reviewer raises the concern that the BMA measurements may capture some trabecular BMD, however it should be noted that the correlation of effect sizes for the BMA/BMD overlapping causal SNPs is negative (-0.95). If our measure of BMA had been erroneously capturing trabecular BMD then we would expect to see a positive correlation of effect sizes for the BMA/BMD overlapping causal SNPs, not a negative one.

      Next, integrating mapped genes with existing bone marrow single-cell RNA-sequencing data revealed patterns of adipogenic lineage differentiation and lipid loading. The problem here is that the scRNAseq studies of the Bone Marrow niche are overwhelmingly mouse. The authors might wish to justify why they are relevant to humans (in the absence of the human-specific scRNAseq).

      We thank the reviewer for pointing this out. We noticed that, although Figure 4 and the Methods do explicitly state that the scRNAseq data is from mouse, it is not stated in the text of the Results. This is now corrected.

      The mouse is commonly used as the model organism for in vivo investigation of human phenotypes and bone marrow adiposity is no exception because, although mice have lower bone marrow adiposity than humans, the timing and sequence in bone marrow adiposity development are similar. BMA research makes extensive use of mouse models literature as exemplified by this review of research within the field (https://www.frontiersin.org/journals/endocrinology/articles/10.3389/fendo.2016.00127/full) and this article recent article (Koh et al. 2024. “Adult skull bone marrow is an expanding and resilient haematopoietic reservoir”. https://www.nature.com/articles/s41586-024-08163-9)

      We updated the results section (line 279):

      “Mesenchymal stem cells of the BM niche commit to either the adipogenic or the osteogenic lineage (Figure 4A) and both the number committing to the adipogenic lineage and their level of lipid-loading influences the total level of BMA. This aspect of BM biology is shared between humans and mice (29), so we made use of an existing mouse scRNAseq dataset of BM mesenchymal lineage cells (30) to study variation in the expression of BMA-associated genes as cells differentiate (Figure 4B).”

      For genetic correlation analysis, the authors selected 7 body and 6 brain traits. The latter traits reflect cognition (general cognitive ability and educational attainment) and brain-related disorders. This selection might seem arbitrary. The interpretation of genetic correlation with cognitive ability, education, and Parkinson's disease was attributed to the recently discovered vascular channels that link calvarial bone marrow to the meninges. This is a fascinating hypothesis, which requires functional proof. However, there might be simpler explanations. Thus, the diploe and the inner table of the calvarium are drained by the same veins as the dura. From the anatomy textbook, we know that diploic veins connect the pericranial and endocranial venous system through the skull.

      Whilst it is true that we did not systematically compare the results of the BMA GWAS to all potentially relevant brain and body phenotypes, we did use criteria to select the phenotypes we compared to. As stated in the manuscript (line 304):

      “We selected body traits (BMD, BMI, waist-to-hip ratio, systolic and diastolic blood pressure, type-2 diabetes, coronary artery disease) that have a logical connection to BMA given the mesenchymal stem cells origin of BM adipocytes and their role in bone, fat, and vasculature (29). For the brain, we selected traits reflecting cognition (general cognitive ability and educational attainment) and disorders that are prevalent in adulthood (insomnia, multiple sclerosis, Parkinson's disease, Alzheimer's disease) since it is primarily in adulthood that the adiposity of calvarial BM experiences a substantial change”

      We entirely agree that the suggestion that the genetic correlation between BMA and cerebral traits may be mediated by the vascular channels linking calvarial bone marrow to the meninges is merely a hypothesis. We have therefore updated the text of the Discussion (line 470):

      “We tentatively speculate that calvarial MALPs may be involved in sensing perivascular flows of CSF from the meninges to the BM and in influencing the BM’s hematopoietic response, and that this might be the basis of the observed genetic overlap between BMA and some cerebral traits. However, more conventional anatomical pathways may also be relevant, as the diploë and inner table communicate with meningeal and dural venous systems through diploic veins.”

      Reviewer #2 (Public review):

      Summary:

      This study develops a new artificial intelligence method for high-throughput analysis of skull bone marrow from MRI data, which may be useful for large-scale biological analyses. Using this method, the authors then attempt to estimate skull bone marrow adiposity (BMA) using T1-weighted signal intensity from MRI scans of ~33,000 people, followed by genome-wide association analysis; however, the approach is inadequate because T1-weighted signal intensity is not validated for measurement of bone marrow adiposity. If it could be validated, the study would be an important advance in understanding of bone marrow adiposity and skeletal biology.

      Strengths:

      This paper is well-written, and the figures are nicely presented. The neural network method used for analysing skull bone marrow is innovative, and the authors validate this through several approaches. Therefore, the authors have achieved the aim of developing a method for large-scale analysis of skull bone marrow from MRI data.

      The GWAS is reasonably well-powered and addresses potential ethnicity differences, with one GWAS done across white males and females, and a separate GWAS in non-white participants. The methodology also conforms to common GWAS standards, including for mapping genetic variants to candidate genes. Moreover, the study further investigates the biological roles of these genes by analysing their expression in single-cell RNA sequencing data.

      Weaknesses:

      The fundamental weakness is that T1-weighted MRI signal intensity (T1W) is used as an estimate of BMA, but it has never been validated for this. The authors show that this T1W parameter measures something that is heritable and can be compared between subjects, but they don't show that it actually measures (or even estimates) calvarial BMA. There is an attempt to do so by comparing the T1W parameter with data from quantitative T1 images: the authors show a reasonable correlation with some of the quantitative T1 image data. However, this still does not show that the parameter is measuring BMA; it could be measuring some other biological characteristic, but this remains unclear. So, there is a need to validate the T1W parameter against an established measure of BMA, such as the bone marrow fat-fraction or proton density fat fraction measured from multi-echo MRI analysis.

      Without validating this BMA measurement method, it is not possible to interpret the GWAS or other findings reported in the study.

      We reject this criticism.

      Although T1-weighted has not been validated as a quantitative measure of fat-fraction, there are several studies showing that it is a semi-quantitative measure of fat content (e.g. Loevner et al 2002, Shen et al 2013, Zhang et al 2020) and we also provide data that support this (figures S9-11).

      Semi-quantitative measures are used in many biomedical GWASes for instance even highly heritable neuropsychiatric disorders (such as schizophrenia and bipolar disorder) involve assessment by clinicians where the test-retest kappas are in the range 0.4-0.6.

      Further, we would suggest that the shortcoming of the imperfect correlation of T1w signal intensity with fat content is more than outweighed by our precision in identifying the calvarial BM cavity and the fact that the flat calvarial bone marrow has a wide range of adiposity in middle-aged and elderly individuals (compared to other bones). This lies at the root of why:

      We clearly recapitulate the known sex and age profiles, as well as the effect of HRT.

      We estimate high BMA heritabilities (43% in males and 23% in females)

      We find clear sex differences (which is a known feature of BMA biology)

      We identify a large number of the genes already known to affect BMA from earlier animal and cell work

      A noisy measurement of an entity with strong biological signal (a well-defined bone marrow cavity with variation in BMA across subjects) will often be more informative than a highly precise measurement of a poorly defined entity with little signal.

      A less critical weakness is that the GWAS has been done only on a single cohort, without replicating the findings in a follow-up cohort. For example, the authors could repeat their analysis on the remaining ~50,000 UK Biobank imaging participants for whom MRI data is now available. However, this would be pointless without knowing what biological characteristic(s) the T1W parameter is actually reflecting.

      We disagree with this comment. We separated the UKB data into discovery and replication sets prior to running the GWAS, so these datasets are independent:

      (1) Further, we ran the discovery (white british individuals) GWAS separately for males and females (prior to combining) and reported in the results section: “We found them to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      (2) We performed our replication GWAS in non-white British males and females. As noted in the results section: “Out of the 168 significant discovery SNPs, 62% replicated at P<.05, and 39% of the 41 lead SNPs replicated at P<.05 (Table S6). One locus replicated at genome-wide significance (P<5e-8). Furthermore, 92.7% of the lead SNPs of the discovery sample showed same effect direction in the replication sample (Table S5).”

      Reviewer #3 (Public review):

      Summary:

      This manuscript, "Estimating bone marrow adiposity from head MRI and identifying its genetic 2 architecture", brings together the groups of Drs. Kaufmann and Hughes in a tour de force work to develop an artificial neural network that localizes calvaria bone marrow in T1-weighted MRI head scans, with the goal of studying its composition in several large MRI datasets, and to model sex-dimorphic age trajectories, including the effect of menopause.

      Strengths:

      Bone marrow adiposity is a very active tissue with far-reaching implications for tissue crosstalk and human health than we had initially recognized. Although MRI has been used to measure BM, studies such as the one by these two groups are still lacking whereas very large datasets are analyzed using advanced AI machine learning tools coupled with genetic studies and a specific pathology. The groups had to develop new methods and new AI machine-learning tools for the imaging analyses.

      Weaknesses:

      Some aspects of the work that authors could add additional clarification.

      (1) Imaging Limitations: The authors provide an excellent overview and references supporting the use of MRI as a method for assessing marrow fat, particularly with some specific modifications. However, MRI images can be affected by various factors, including the presence of other tissues as well as specific MRI settings, which are much harder to precisely control when using different datasets.

      We thank the reviewer for his positive assessment of our review of methods.

      Regarding MRI settings: We agree with the reviewer that differences in scan protocols can create substantial differences in the resulting images between samples. Different tools exist for harmonization of imaging data across sites, but they usually operate on tabulated data and there is no one-size-fits-all approach yet [1]. Here, we took a different approach to prevent confounding bias: We generated a large set of simulated data for training of the neural network. The simulations circumvented potential issues emerging from confound biases in training sets that we might have seen had we had combined multiple samples with different scan protocols. Nevertheless, applied to real data the models may still face confound issues, such as better BMA estimates for some scan protocols over others. We have addressed these issues as follows: (1) Validation analyses (10 repeat scans of 30 individuals and twin pairs, figure 1d and 1e) are based fully on data that was acquired on the same scanner with the same protocol. (2) Analysis in UK Biobank included data from different scan sites albeit harmonized protocols. Here we accounted for scan site in all statistical models (including GWAS).

      (1) Dominik Kraft, Gloria Matte Bon, Édith Breton, Philipp Seidel, Tobias Kaufmann; Removing scanner effects with a multivariate latent approach: A RELIEF for the ABCD imaging data?. Imaging Neuroscience 2024; 2 1–7. doi: https://doi.org/10.1162/imag_a_00157

      Regarding the presence of other tissues: We recognise in the existing text of the results section that sometimes inner or outer table voxels are wrongly identified as bone marrow, but we show that this does not have a major impact on the correct identification of the bone marrow cavity. The existing text reads:

      “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap. However, this did not result in a corresponding fall in the ratio of the predicted intensity of BM to its true intensity, because the typical BM intensity was only marginally higher than the neighbouring bone intensity. This property of the typical relative intensities of these anatomic structures also explains why the intensity ratio at high overlap is not centred on 1: any misidentification of cortical bone as BM, will typically result in an underestimate of true BM intensity (Figure S2). The neural network performed well and intensity ratios were in the range 0.9-1.1 for the vast majority of head types (Figure S3).”

      Also note that we implement a number of QC measures to exclude scans where there is evidence that we may have failed to correctly identify the bone marrow cavity. The existing text reads:

      “We used two additional QC metrics to filter out calvaria where BM location was likely to have failed. First, we set an upper limit of 30 on the standard deviation of the intensity of the outer table as scans with higher values were clear outliers and were probably cases where the location of both outer table and BM has failed (Figure S6). Second, for each calvarium, we computed the Mahalonobis distance for all vertices in the two dimensions “first layer of the BM” and “BM intensity” (Figure S7). By manual inspection we found that data points with MD > 25 often had errors in BM layer identification, typically where the network had erroneously predicted a higher and more intense layer to be the BM. We considered a calvarium as failing this QC criterium if more than 0.5% of vertices have MD > 25. This criterium is very strict as errors on only 0.5% of data points in a calvarium would not significantly affect the average BM intensity for a calvarium.”

      (2) The specific density of cranial bones as it relates to the types of bone marrow: Cranial bones are extremely dense structures, which naturally interfere with MRI imaging. While it is thought that cranial bones have mostly "red bone marrow", this is only true for a short time in humans. How sensitive is their system in differentiating between red and yellow BM?

      We implemented several measures to ensure that our method would be robust to anatomical variation between individuals. As noted in the current version of the Methods section: “In order to train the neural network model, we generated a large synthetic dataset of intensity arrays, with known boundaries between anatomical structures, by simulating the thickness and intensity of the different structures located between the outer skin and the subarachnoid space. The simulation incorporated the following real-world complexities:

      Different anatomical architectures (skin, subcutaneous fat, aponeurosis, outer table, BM, inner table, dura mater, arachnoid space), including when a structure is not present throughout the calvarium

      Variation in thickness and intensity between vertices (on the same calvarium)

      A wide variety of different calvarium types with different combinations of levels of BM adiposity, bone thickness, and subcutaneous adiposity.”

      Further, as noted in the Results section:

      “We evaluated the performance of the neural network on simulated data using two metrics (Figure 1B): 1. the overlap between the predicted and the true BM location, and 2. the ratio between the predicted intensity of the BM and the true intensity of the BM”. The accuracy in localising the bone marrow layers was good, with the only exception being: “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap”

      We also validated our procedure on real data (see Results section, subsection “Procedure validation on real data and heritability estimate”). Briefly, we checked the accuracy of our method using a dataset from the Consortium for Reliability and Reproducibility, a twin dataset from the Human Connectome Project and by manually checking many hundreds of UKBiobank scans.

      We are thus confident that we accurately identify the bone marrow cavity irrespective of whether the bone marrow is red (low adiposity) or yellow (high adiposity).

      (3) Both items above are further complicated by aging, but aging is not a linear event as we have learned. There are specific bursts of aging in humans around the age of 45 and early 60s. How do the system and model predict or incorporate these peaks of aging? It seems from the data shown that aging is reflected more as a linear phenomenon. Is this because additional aging datasets are needed?

      We agree with the reviewer that ageing probably occurs in bursts rather than being a linear process. We do see a non-linear relationship between age and BMA in our data (see figure 2B), with a more rapid rise in BMA between the ages of 45 and 65, than later in life (in women). As a result of this, when we model BMA using regression, we use orthogonal polynomials of degree 2 which allows for a non-linear relationship. However, we cannot observe bursts of BMA increase in our data because it is cross-sectional. Longitudinal data would be required to obtain information on the nature and timing of any bursts in bone marrow adiposity.

      (4) The authors describe in richness of detail their AI learning programming and how it extracted the data from datasets. The authors also show some important correlations with specific genes, SNPs. What is not clear is how conditions such as anemia for example. An expected finding would be that patients with chronic anemia have lower bone marrow (BM) signal intensity on MRI scans than healthy people. This is because the signal intensity of BM depends on the fat-to-cell ratio in the tissue.

      We agree with the reviewer that conditions affecting the bone marrow niche have a potential to affect and be affected by bone marrow adiposity, with leukemia being a known example. This is why we believe that a method, such as the one we present here, has the potential to be useful in several biomedical fields (hematology and osteology).

      Furthermore, patients with a host of musculoskeletal disorders ranging from osteopenia to osteoporosis, sarcopenia, and osteosarcopenia will also have altered MRI scans. When using such large datasets how did the authors control or exclude these pathological conditions, or were all these conditions likely present?

      We did not exclude specific pathologies. We were careful to train our NN model on a wide variety of skull thicknesses, bone marrow adiposity, and subcutaneous adiposity and to evaluate the performance of the model on simulated and real datasets (see answer to your point 2).

      Reviewer 1 raised the issue of individuals displaying Hyperostosis frontalis interna (thickening of the inner table) and we recognize that in extreme case of this condition, where there is a major change in the anatomy of the calvarial bone, our method would probably not correctly localise the bone marrow. However, such extreme cases are rare and thus would not have a major impact on our results derived from over thirty thousand individuals. We demonstrate this with an extra analysis performed in response to the point about hyperostosis frontalis interna made by reviewer 1.

      (5) Some of the genes and SNPs although significant showed very small correlations. What is their likely physiological significance?

      Bone marrow adiposity is a polygenic trait and we have identified 41 statistically significant loci. We had a discovery sample of approximately 30k individuals which is modest for a GWAS study, so these 41 loci are a lower bound on the number of genes influencing the BMA trait. When a large number of genes influence a trait, the effect size of an individual gene is typically relatively small. However, the SNP heritability estimates of 31.5% indicates that we are able to explain approximately one third of the phenotypic variation with the effect sizes estimated by our GWAS: this is quite a high fraction relative to many other GWASs of biomedical traits.

      (6) The authors could use this excellent manuscript to expand their discussion to include the need for studies like theirs to be also complemented by multi-OMICS studies that will include proteomics and lipidomics of BM, bones, and muscles.

      We agree with the reviewer and hope that such studies will be undertaken in the future. We attempted to point in this direction in the last sentence of the Discussion (line 517): “Future studies can build on our developments to further validate the proposed measure of bone marrow composition and to study its effect on bone, blood, and brain”. Word count limits prevented us from further expanding on the specific kinds of studies that should be performed.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      More moderate concerns include:

      (1) In the "Genetic correlation and overlap" part, it is unclear why a high effect correlation with high SNP overlap is suggestive of vertical pleiotropy, while "moderate overlap and a low correlation of effect sizes ... are more indicative of horizontal pleiotropy". This is not intuitive.

      In vertical pleiotropy (genetics > phenotype A > phenotype B): genetics drive phenotype A and phenotype A drives phenotype B (B has few direct genetic drivers of its own). If we perform a GWAS of phenotype A and a GWAS of phenotype B, one would expect to see a high overlap in the causal variants (because phenotype B is largely indirectly determined by the genetics of phenotype A) and the correlation should be high (either negative or positive) because there is a cause-effect relationship between A and B (whether phenotype A has a positive or negative effect on phenotype B).

      In horizontal pleiotropy: the A and B phenotypes share the same genetic loci (but have no phenotypic influence on each other). In this case, we would observe high overlap in the associated loci, but one would not expect to see a high correlation in effect sizes (across loci) because there is no a priori reason to expect that genes associated with both phenotype A and phenotype B would have a consistent (negative or positive) effect across loci. For example, gene X may increase A and B, whereas gene Y may increase A but decrease B.

      We have made a small update to the relevant part of the Results section and otherwise rely on the explanation above (which will be publicly available along with the manuscript):

      Line 325: “These patterns of high overlap and high effect correlation within this overlap are suggestive of vertical pleiotropy i.e. a molecular mechanism influencing one trait, that in turn influences a second trait, such that most of the variants driving the first trait either have the same or the opposite direction of effect on the second trait.”

      (2) "possible causal effect of BMA on cognition" asks for a formal analysis, like Mendelian randomization.

      Given the current wording, the reviewer is justified in asking for a formal analysis. Since we did not perform this analysis, we have changed the wording:

      Line 434: “Since these are two highly correlated traits [36], a high overlap and correlation of genetic effects for both traits with BMA may be consistent with the hypothesis that BMA could have a causal effect on cognition (Table 1).”

      (3) ll. 132-134: please reword this sentence for clarity: "Using the network-estimated location of the BM within ... averaged these across all datapoints...". Please define threshold of desirable overlap between the predicted BM and the true BM (=0.7?).

      Background: The model predicts the BM localisation for a datapoint (which interval of layers of the 50 layers is bone marrow). We tested the model on a wide variety of simulated data and aim for the overlap to be as close to 1 as possible, but some error is inevitable. We found that the average overlap between the true and predicted bone marrow was only below 0.7 when the bone layer is only 4 mm thick (meaning that the bone marrow is only 1-2 mm thick). This demonstrates the high accuracy of our method in identifying a very small anatomical feature.

      When applying our method to real data, we do not know the truth and therefore cannot compute the overlap between the predicted and true value. It is therefore not possible to identify datapoints where the overlap is poor (e.g. lower than 0.7) and filter them out.

      Given the above, we struggle to understand in what way an overlap threshold is relevant to how we compute the signal intensity for a datapoint. Nevertheless, we recognize that the sentence pointed to by the reviewer is poorly formulated and have tried to make it clearer:

      Line 130-133: “To obtain the BM signal intensity for an individual datapoint of the calvarium, we used the network model to estimate the location of the BM within the datapoint’s intensity array and averaged these BM intensities to get the BM intensity for that datapoint. Then, we averaged these datapoint intensities across the calvarium to produce the global BMA measure for the scan.”

      In the GWAS Results, please clarify the phrases - what was "significantly genetically correlated (Rg=.94)" (also, l. 402, "genetic correlation between the sexes" - in what?).

      Genetic correlation is a statistical measure that quantifies the extent to which two traits (or the same trait in two different cohorts) are influenced by the same genetic factors. Simply put, it is the effect sizes of the SNPs in the two GWASs of interest that are correlated (after correcting for confounding effects, such as linkage desequilibrium). When comparing two GWASs, the standard formulation is to refer to their “genetic correlation”. We made a modification to the text to clarify this:

      Line 239: “We found the male and female GWASs to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      "a more than two-fold difference between the sexes" - in which metric?

      We feel that what is being compared is stated clearly in the original sentence:

      Line 262: “A comparison of the male and female effect sizes of the top lead SNPs of each locus revealed 6 loci in which there is a more than two-fold difference between the sexes (loci 10, 18, 26, 30, 32, 37 in Table S5)”.

      (4) Also In GWAS Results, a locus Dlx5 is called "SHFM" in the Supplementary Table.

      Background:

      We identified 41 genome-wide significant loci and named the locus after the gene closest to the top lead SNP (bold in Figure 3C). Other genes in each locus for which genome-wide significant SNPs were eQTLs, are listed below the closest gene in normal font (Figure 3C).

      In table S5, we report details of the top lead SNP for all 41 loci. We report only the nearest gene to the top lead SNP.

      For locus 14, SHFM1 is the closest gene to the top lead SNP whereas DLX5 and DLX6 are genes in the locus for which genome-wide significant SNPs were eQTLs. This explains why DLX5 appears under SHFM1 in Figure 3C, but does not appear in Table S5.

      (5) Please reword MRI jargon - "Dixon method", vertix - should be introduced, as well as abbreviation "KDE".

      We had recognised that the word “vertex” would be confusing and had replaced it by datapoint, but had unfortunately missed one occurrence in the text. This is now corrected.

      Thank you for pointing out the lack of introduction of the term “KDE”. This was only explained in the supplementary materials, but has now been added to the main text:

      Line 500-508: “To ensure between-subject comparability, we used the intensity normalised nu.mgz volume output by FreeSurfer. We validated this approach through comparison with well-established intensity normalization methods; Kernel Density Estimation (KDE), WhiteStripe (WS), Gaussian Mixture Model (GMM), Fuzzy C-Means (FCM), and Z-score normalization (ZS). We found the highest test-retest reliability with our approach (Figure S9), and, together with KDE (based on reference signal intensity in WM), the highest correlation with quantitative T1 relaxation maps (Figure S10).”

      Reviewer #2 (Recommendations for the authors):

      (1) This would be an extremely useful advance for the bone and BMA fields if only it could be confirmed that the T1W signal intensity is actually measuring BMA in some meaningful way. Or, even if not BMA, to confirm what other biological characteristic(s) it is in fact capturing. This is essential for interpreting the findings.

      (2) I note that you have compared the normalized T1W parameter with quantitative T1 data (e.g. Figure S10). However, this doesn't address the fundamental issue, because even these quantitative T1 data (e.g. from MP2RAGE) may not be measuring calvarial BMA. T1W sequences have been used to estimate BM cellularity (if not BMA directly) but are not nearly as precise as water-fat imaging. For example, one study found a reasonable correlation (0.71) between T1 relaxation times and BM fat (https://www.nature.com/articles/s41598-019-57030-5). So, if your normalized T1W parameter shows a correlation of -0.44 with the T1 MP2RAGE MRI signal (Figure S10), what does this mean in terms of how well your parameter reflects the actual BMA adiposity? We can't know this, because we also don't know if the T1 MP2RAGE signal reflects calvarial BMA.

      (3) I think my recommendations are clear from the public review. Ideally, you would be able to compare the skull BM normalized T1W parameter with PDFF data that have T2* correction (since the skull BM cavity is quite small and so may suffer from T2* effects relating to tissue inhomogeneity). But even if you had only dual-echo BMFF data, this would still be much more informative than relying only on T1 data. I hope this can be done so that the findings of the study can be properly interpreted.

      As explained above, we reject this reviewer’s claim that T1-weighted signal intensity cannot be used to perform a GWAS of BMA: other studies have shown that T1-weighted signal intensity is a semi-quantitative measure of fat fraction, we have performed extra analyses that confirm this, and our results further demonstrate this. For further detail on why we reject this criticism, see our response to this reviewer’s comments.

    1. eLife Assessment

      This work presents valuable new data on the role of D-Serine and how it competes with its stereoisomer L-Serine to influence metabolism. The work presents a variety of convincing experimental data combined with simulated results to investigate the mechanisms focused on one-carbon metabolism, which is relevant for several research fields. However, some claims are only partially supported by data, and critical areas comparing L- vs D-Serine and further mechanistic studies are required. Furthermore, while the work has potential for various fields, the work has only been studied in a limited cell type and context.

    2. Reviewer #1 (Public review):

      Summary:

      The authors demonstrate the stereoselective role of D-serine in 1C metabolism showing that D-serine competes with L-serine and inhibits mitochondrial L-serine transport. They observe expression of 1C metabolites in their metabolomics approach in primary cortical neurons treated with L-serine, D-serine and mixture of both. Their conclusions are based on the reduction in levels of glycine, polyamines and their intermediates and formate. Single cell RNA sequencing of N2a cells showed that cells treated with D-serine enhanced expression of genes associated with mitochondrial functions such as respiratory chain complex assembly and mitochondrial functions with downregulation of genes related to amino acid transport, cellular growth and neuron projection extension. Their work demonstrates that D-serine inhibits tumor cell proliferation and induces apoptosis in neural progenitor cells highlighting the importance of D-serine in neurodevelopment.

      Strengths:

      D-amino acids do not merely function as ligands at receptors but have underlying roles in signaling and metabolism. These roles are just beginning to be uncovered. The authors elucidate the metabolic role of D-serine in the context of neuronal maturation by its suppression of mitochondrial L-serine availability for SHMT2 and 1C flux. This is the strength of the manuscript. The implications for the metabolic role of D-serine in neurons is a highlight and underlines its roles in neuronal metabolism.

      Weaknesses:

      These are some minor issues that come up on critical assessment of the manuscript and is only intended to strengthen the manuscript. The comments below are based on the revisions made by the authors including the justification of their approach and rebuttal.

      (1) Kinetic assessment of D-serine versus L-serine: The authors have made reference to prior work by Miyamoto et al. and justify their rationale. This is acceptable.

      (2) Molecular Dynamics simulations while a good first step in modeling interactions at the active site, relies on force fields. The authors state that any elaborate study into longer simulations is beyond the scope and their simulations data are supported by other experimental work. This is justified.

      (3) The use of N2a cell line is also justified to reflect the proliferative nature of immature neurons.

      (4) With regards to caspase 3 comment, the whole blot is convincing and shows cleaved caspase-3 band at approx. 15 kDa.

      (5) Scale Bars are clearly visible and Fig S6 which was earlier S5 is legible. If possible, the authors can include an magnified inset in the merged image to show the clear activation of caspase-3.

      (6) Issue of phosphatidyl serine standard in LC-MS is justified by the use of L-serine standard due to lack of availability.

      (7) The authors mention about enantiomeric shift of serine metabolism during neural development which appears to be a discussion of prior published data from Hubbard et al 2013, Burk et al 2020, and Bella et al 2021 in Supplementary Figure panels 8 A-E.<br /> The authors justify by citing references to the work which may be acceptable and also the current norms of publication. This reviewer felt contrary to the fact, however it is left to the editors to make a decision on this.

      (8) The discussion section has been substantially revised and now reads well.

      (9) The relevant references have been cited. In doing so, the work integrates and elucidates a mechanistic and functional role of D-serine in neurons.

      (10) Figure S7A in the revised manuscript shows the specificity of D-serine in the cleaved caspase-3 assay which is informative.

      Comments on revised version.

      This reviewer is satisfied by the effort made by the authors based on the prior comments raised.

    3. Reviewer #2 (Public review):

      Summary:

      This study by Suzuki et al. reports an interesting stereo-selective role of D-serine in regulating one-carbon metabolism during neurodevelopment to adapt the functional transition, probably through the competition with mitochondrial transport of L-serine. The authors provide a multi-layered set of evidence, including metabolomics, enzyme assays, mitochondrial transport competition and functional assays in immature/neural progenitor cells, to build up a conceptual integration of D-serine as both a neurotransmitter and a metabolic regulator in central neural system, which raises a broad potential interest to the neuroscience and metabolism communities.

      Strengths:

      This work provides a conceptual advance that D-serine is not only serves as a traditional neurotransmitter in central neural system but also critically contributes to metabolic regulation of neural cells. The authors performed solid metabolomic assays to validate the suppressive effect of D-serine on one-carbon metabolic pathway, providing some evidence that D-serine competitively inhibits mitochondrial serine transport, but not directly impairs SHMT2 enzymatic activity. All these data indicate a critical role of D-serine synthesis during neural maturation and suggest a potential translational strategy for targeting serine metabolism in neural tumors.

      Comments on revised version.

      My previous concerns have been appropriately addressed or discussed in this revised version of manuscript. I have to say that, at this stage, I have no further questions.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript presents a comprehensive and well-executed investigation into the metabolic role of D-serine in the central nervous system. The authors provide solid evidence that D-serine competitively inhibits mitochondrial L-serine transport, thereby impairing one-carbon metabolism. This stereoselective mechanism reduces glycine and formate production, suppresses cellular proliferation, and induces apoptosis in immature neural cells and glioblastoma stem cells. Developmental analyses further reveal a physiological enantiomeric shift in serine metabolism during neurogenesis, aligning with the transition from proliferation to maturation. Overall, the study bridges developmental neurobiology, cancer metabolism, and amino acid transport, uncovering a previously unrecognized metabolic function of D-serine beyond its role in neurotransmission.

      Strengths:

      (1) The discovery that D-serine inhibits one-carbon metabolism by competing for mitochondrial L-serine transport-rather than through enzymatic inhibition or receptor-mediated signaling-represents a significant and previously underappreciated mechanism. This finding has broad implications for understanding metabolic regulation during neurodevelopment and offers potential relevance for targeting metabolic vulnerabilities in cancer.

      (2) The authors integrate metabolomics, mitochondrial transport assays, molecular dynamics simulations, genetic and pharmacologic perturbations, transcriptomics, and both in vitro and ex vivo models. The breadth of experimental approaches, combined with the coherence of the findings across systems, provides strong support for the central conclusions and enhances the overall impact of the study.

      (3) The temporal shift in D-/L-serine levels during neurodevelopment is elegantly linked to the transition from proliferative to mature neuronal states. The selective vulnerability of neural progenitors and tumor cells-contrasted with the resistance of mature neurons-highlights a biologically meaningful and potentially targetable metabolic distinction.

      Weaknesses:

      (1) While the authors attribute D-serine's metabolic effects to competition with mitochondrial L-serine transport, the specific identity of the transporter(s) mediating this process remains undefined. This represents a meaningful mechanistic gap, as the central conclusion depends on D-serine limiting mitochondrial L-serine availability to inhibit one-carbon metabolism.

      (2) The effective concentrations of D-serine used in vitro (IC₅₀ ≈ 1-2 mM) exceed typical brain levels (~0.3 mM). While the authors acknowledge this, a more focused discussion on whether higher local D-serine concentrations could arise in specific microenvironments-such as synaptic compartments, tumor niches, or pathological states-would help contextualize the in vitro findings and strengthen their physiological relevance. For example, disruptions in D-serine clearance or altered expression of serine racemase and transporters in disease contexts could lead to localized accumulation. Moreover, differences between extracellular and intracellular D-serine pools-and the mechanisms governing their regulation-may further influence its metabolic impact in vivo.

      (3) While the manuscript focuses on neural stem/progenitor cells and neural tumors, it remains unclear whether the anti-proliferative effects of D-serine are specific to neural lineages or extend to other highly proliferative non-neural cell types. A brief discussion addressing this point would help clarify the scope of D-serine's metabolic impact and whether its mechanism of action reflects a unique vulnerability in neural cells or a more general feature of proliferative metabolism. This distinction is particularly relevant for assessing the broader therapeutic potential of targeting mitochondrial L-serine transport.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors demonstrate the stereoselective role of D-serine in 1C metabolism, showing that D-serine competes with L-serine and inhibits mitochondrial L-serine transport. They observe expression of 1C metabolites in their metabolomics approach in primary cortical neurons treated with L-serine, D-serine, and a mixture of both. Their conclusions are based on the reduction in levels of glycine, polyamines, and their intermediates and formate. Single-cell RNA sequencing of N2a cells showed that cells treated with D-serine enhanced expression of genes associated with mitochondrial functions, such as respiratory chain complex assembly, and mitochondrial functions, with downregulation of genes related to amino acid transport, cellular growth, and neuron projection extension. Their work demonstrates that D-serine inhibits tumor cell proliferation and induces apoptosis in neural progenitor cells, highlighting the importance of D-serine in neurodevelopment.

      Strengths:

      D-amino acids are a marvel of nature. It is fascinating that nature decided to make two versions of the same molecule, in this case, an amino acid. While the L-stereoisomer plays well-known roles in biology, the D-stereoisomer seems to function in obscurity. Research into these novel signaling molecules is gathering momentum, with newer stereoisomers being discovered. D-serine has been the most well-studied among the different stereoisomers, and we still continue to learn about this novel neurotransmitter. The roles of these molecules in the context of metabolism is not well studied. The authors aim to elucidate the metabolic role of D-serine in the context of neuronal maturation with implications for 1C metabolism and in cell proliferation. The metabolic role of these molecules is just beginning to be uncovered, especially in the context of mammalian biology. This is the strength of the manuscript. The authors have done important work in prior publications elucidating the role of D-amino acids. The advancement of the field of D-amino acids in mammalian biology is significant, as not much is known. The presentation of RNA seq data is a valuable resource to the community, however, with caveats as mentioned below.

      Weaknesses:

      The following are some of the issues that come out in a critical reading of the manuscript. Addressing these would only strengthen and clarify the work.

      (1) Kinetic assessment of D-serine versus L-serine: While the authors mention that D-serine is not a good substrate for SHMT2 compared to L-serine, the kinetic data are presented for only D-serine. In a substrate comparison with an enzyme, data must be presented for L-serine as well to make the conclusion about substrate specificity and affinity. Since the authors talk about one versus another substrate, there needs to be a kinetic comparison of both with Km (affinity). (Ref Figure 2 panel).

      We agree with the reviewer that kinetic parameters for l-serine are important for evaluating the substrate specificity of SHMT2. The hydroxymethyl-transferase activity of SHMT2 toward l-serine has been previously characterized by our co-author Tetsuya Miyamoto (Miyamoto et al., FEBS Journal, 2024; PMID: 37700610), which is appropriately cited in the manuscript (line 143). In that study, the kinetic parameters for l-serine were determined, with a Km of 0.07 ± 0.009 mM and a Kcat of 33.5 ± 0.9 min<sup>-1</sup>. The strong chiral selectivity of SHMT2 for the l-enantiomer in the hydroxymethyl-transferase reaction was also demonstrated in that work. On the other hand, the primary aim of the present study is different from characterizing d-serine as a catalytic substrate for SHMT2. Rather, our goal was to determine whether d-serine interferes with the hydroxymethyl-transferase reaction of SHMT2 by interacting with the l-serine binding site. Accordingly, the analyses shown in Fig. 2B-D were designed to evaluate whether d-serine could structurally occupy or interfere with the l-serine binding pocket of SHMT2. Therefore, our experiments focused on assessing the potential inhibitory effect of d-serine rather than performing a full kinetic comparison of d-serine and l-serine as substrates.

      (2) Molecular Dynamics simulations, while a good first step in modeling interactions at the active site, rely on force fields. These force fields are approximations and do not represent all interactions occurring in the natural world. Setting up the initial conditions in the simulations can impact the final results in non-equilibrium scenarios. The basic question here is this: Is the simulated trajectory long enough so that the system reaches thermodynamic equilibrium and the measured properties converge? Prior studies have shown mixed results with the conclusion that properties of biological systems tend to converge in multi-second trajectories (not nanosecond scales as reported by the authors) and transition rates to low probability conformations require more time. (Ref Figure 2C).

      We thank the reviewer for raising the important point regarding the limitations of molecular dynamics (MD) simulations, including the dependence on force fields and the potential effects of simulation length and initial conditions. We agree that MD simulations represent approximations of molecular behavior and that longer trajectories may be required to fully explore rare conformational states in biological systems.

      In the present study, however, the MD simulations were not intended to provide a comprehensive thermodynamic description of SHMT2 conformational dynamics. Rather, they were used as a structural assessment to evaluate whether d-serine could plausibly occupy the canonical l-serine binding site of SHMT2. As shown in Fig. 2BC and supplementary movie 1, the simulations did not support stable occupation of the l-serine binding pocket by d-serine. Importantly, this structural observation is consistent with our biochemical data showing that d-serine does not inhibit the hydroxymethyl-transferase activity of SHMT2 when l-serine is used as the substrate (Fig. 2D). Together, these results indicate that the inhibitory effect of d-serine on one-carbon metabolism is unlikely to be mediated through direct inhibition of SHMT2.

      As the reviewer correctly notes, it remains possible that d-serine interacts with SHMT2 at sites distinct from the canonical l-serine binding pocket and could exert potential allosteric effects. Indeed, previous work by Miyamoto et al. (FEBS Journal, 2024) demonstrated that SHMT2 exhibits dehydratase activity toward d-serine. However, this reaction is not directly linked to mitochondrial one-carbon metabolism. Therefore, further extensive simulations exploring alternative conformational states or potential allosteric interactions would extend beyond the scope of the present study.

      Importantly, the key conclusions of this study do not rely solely on MD simulations but are supported by multiple independent experimental approaches, including metabolomics, enzymatic assays, and mitochondrial transport analyses.

      (3) The authors use N2a cell line to demonstrate D-serine burden on primary cortical neurons. N2a is an immortalized cell line, and its properties are very different from primary neurons. The authors need to mention a rationale for the use of an immortalized cell line versus primary neurons. The transcriptomic profile of an immortalized cell line is different compared to a primary cell. Hence, the response to D-serine may vary between the two different cell types.

      We thank the reviewer for raising this important point regarding the differences between immortalized cell lines and primary neurons. As the reviewer notes, N2a cells and primary cortical neurons (PCNs) differ in several aspects, including their degree of differentiation and proliferative capacity. We appreciate the opportunity to clarify the rationale for using both systems in this study.

      One-carbon metabolism is known to be particularly active in highly proliferative or relatively undifferentiated cells. Primary cortical neurons are initially obtained as immature neuronal populations and gradually undergo maturation during culture (Fig. S6). In our experiments, we observed that sensitivity to d-serine and dependence on one-carbon metabolism were primarily evident in immature neuronal states rather than in fully mature neurons (Fig. 4EF).

      In this context, immature PCNs share certain metabolic characteristics with proliferative neural cell lines such as N2a cells. Consistent with this idea, the inhibitory effects of d-serine on one-carbon metabolism and cell proliferation were observed in both immature PCNs and N2a cells (Fig. 1E–G, Fig. 2E, Fig. 3AB, and Fig. 4F). Thus, the use of N2a cells provides a complementary experimental model for studying the metabolic vulnerability of immature neural cells that depend on one-carbon metabolism. Importantly, the key findings were consistently reproduced in primary cortical neurons, supporting the physiological relevance of the observations made in N2a cells. For clarity, we added descriptions in lines 107-108 and 167-168 in our revised manuscript.

      (4) In Figure 4D, the authors mention that D-serine activates the cleavage of caspase 3. Figure 4D shows only cleaved caspase 3 as a single band. They need to show the full blot that contains the cleaved fragments along with the major caspase 3 band.

      In our experiments, we used an antibody that specifically recognizes cleaved caspase-3 and does not recognize full-length caspase-3 (Cell Signaling Technology, anti-cleaved caspase-3 antibody, clone 5A1E). Therefore, the Western blot detects only the cleaved caspase-3 fragment at approximately 17 kDa, which appears as a single band in the blot (please see Author response image 1). For clarity, we have also included the antibody information in the Western blot section in the revised Materials and Methods (lines 501-502).

      Author response image 1.

      An original image of western blot for Fig. 4D. An arrow indicates the bands of cleaved caspase-3 (17 kDa).

      (5) In Figure panel 4, the authors use neural progenitor cells (NPCs). They need to demonstrate that the population they are working with is NPCs and not primary neurons. There must be a figure panel staining for NPC markers like SOX2 and PAX6. Also, Figure S5 needs to be properly labeled. It is confusing from the legend what panels B-E refer to? Also, scale bars are not indicated.

      We thank the reviewer for pointing out these issues. First, we have relabeled and rearranged the figure panels in revised Figure S6B-E, and revised the figure legend to provide a clearer description of each panel. We have also added scale bars to the revised images.

      To confirm the identity of neural progenitor cells (NPCs), we performed immunostaining for Nestin, a well-established marker for NPCs, instead of Sox2 and Pax6 (Fig. S6B). Nestin is widely used as a marker for NPCs(Bernal and Arranz, 2018; Bott et al., 2019; Lendahl et al., 1990), and is known to be co-expressed with Sox2 in the mouse embryonic brain(Graham et al., 2003). In addition, Nestin has been identified as a Pax6-bound gene associated with the transcriptional program regulating neural progenitor identity (Thakurela et al., 2016). Consistent with these observations, analysis of published scRNA-seq data from the developing mouse brain (Bella et al., 2021) shows that the RNA expression profile of Nestin (Nes) closely parallels those of Sox2 and Pax6 (new Fig. S9C). Based on these lines of evidence, we consider Nestin-positive cells in our cultures to represent NPCs. Importantly, Nestin-positive cells were co-stained with cleaved caspase-3, suggesting that the apoptotic population corresponds to NPCs (Fig. S6B). Together, these observations support that the apoptotic cells observe in our culture correspond to NPCs rather than differentiated neurons.

      (6) In Supplementary Figure panel 7F, the authors mention phosphatidyl L-serine and phosphatidyl D-serine. A chromatogram of the two species would clarify their presence as they used 2D-HPLC. On an MS platform, these 2 species are not distinguishable. Including a chromatogram of the 2 species would be helpful to the readers.

      We thank the reviewer for this helpful suggestion. As the reviewer correctly noted, mass spectrometry alone cannot distinguish phosphatidyl-d-serine and phosphatidyl-l-serine. To quantify these species separately, lipids were first extracted from cells using the Bligh and Dyer method, followed by phospholipase D treatment to cleave the serine moiety from phosphatidylserine (a schematic of the procedure is shown in Fig. S8D). The released serine was then derivatized with NBD-F and analyzed by 2D-HPLC for enantioselective separation and quantification of d- and l-serine, as described in the Materials and Methods section (“Quantification of glycine and serine enantiomers”). Phosphatidyl-l-serine (Sigma-Aldrich: P0474) was used to generate a standard curve for quantification (Fig. S8E). Because phosphatidyl-d-serine is not commercially available, and because the peak height of free d-serine is equivalent to that of free l-serine in the chromatograms of our 2D-HPLC system, the same standard curve was used to estimate phosphatidyl-d-serine levels.

      As requested by the reviewer, we have now added representative chromatograms of d- and l-serine derived from phosphatidyl-serine in NPC samples (new Fig. S8F).

      (7) The authors mention about enantiomeric shift of serine metabolism during neural development, which appears to be a discussion of prior published data from Hubbard et al, 2013, Burk et al, 2020, and Bella et al, 2021, in Supplementary Figure panels 8 A-E. This should not be presented as a figure panel, as it gives the false impression that the authors have performed the experiment, which is clearly not the case. However, its discussion can well serve as part of the manuscript in the discussion section.

      We appreciate this helpful suggestion. The datasets used in Fig. 4K and Fig. S9 were derived from previously published transcriptomic studies (Hubbard et al, 2013; Burk et al, 2020; Bella et al, 2021). While these studies reported transcriptomic profiles during neuronal development in vitro or in vivo, they did not specifically analyze serine metabolism, one-carbon metabolism, or d-serine biosynthesis, which are the focus of the present study. Therefore, we reanalyzed these publicly available datasets from a metabolic perspective, focusing on genes involved in serine metabolism and one-carbon metabolism. This re-analysis allowed us to examine a developmental shift in serine enantiomer metabolism, we presented the results as figure panels rather than simply citing the datasets in the Discussion. Re-analysis of publicly available transcriptomic datasets to address new biological questions has become a common approach in genomics and transcriptomics studies.

      To avoid the impression that these experiments were performed in this study, we have clearly indicated the original references and clarified this point in the figure legends (Fig. 4K and Fig. S9). These panels are intended to provide supportive evidence for the developmental shift in serine enantiomer metabolism discussed in this study.

      (8) The entire presentation of the section on enantiomeric shift of serine metabolism during neural development (lines 274-312) is a discussion and should be part of the discussion section and not in the results section. This is misleading.

      We thank the reviewer for this comment. This point overlaps with the concern raised in comment 7. Please see our response to comment 7 for a detailed explanation and the revisions made in the manuscript.

      (9) The discussion section is not well written. There is no mention of recent work related to D-serine that has a direct bearing on its metabolic properties. In the discussion section, paragraph 1, the authors mention that their work demonstrates the selective synthesis of D-serine in mature neurons as opposed to neural progenitor cells. This concept has been referred to in prior publications:

      (a) Spatiotemporal relationships among D-serine, serine racemase, and D-amino acid oxidase during mouse postnatal development. PMID:14531937.

      (b) D-cysteine is an endogenous regulator of neural progenitor cell dynamics in the mammalian brain. PMID:34556581.

      We thank the reviewer for this helpful suggestion and for drawing our attention to these studies. We have revised the Discussion to better place our findings in the context of previous work related to d-serine.

      Specifically, we have added references describing the spatiotemporal relationship between d-serine and serine racemase (Srr) during brain development (PMIDs 14531937 and 33592203) (line 333). These studies highlighted the tissue-level (PMID 14531937) and cellular-level (PMID: 33592203) relationship between Srr expression and development. These studies highlighted the role of d-serine in supporting the functional maturation of neurons during postnatal development. In contrast, our study addresses a complementary question of why d-serine is NOT present during embryonic and early postnatal stages of brain development, when proliferative metabolic activity is high. This question is fundamentally different from the previous reports. Our point is that d-serine is not favorable because it interferes with one-carbon metabolism, which is essential for cell proliferation. Therefore, this concept has not been referred to in prior publications. To clarify this point, we added the following sentence to the Discussion (line 325-328): “In addition to the known role of d-serine in the functional maturation of differentiated neurons, our findings highlight a previously unrecognized, stereoselective regulation of cellular metabolism by d-serine, and provide a rationale for its selective synthesis in mature neurons where proliferative metabolic activity is no longer required’.

      We also appreciate the reviewer bringing our attention to the study describing d-cysteine as a regulator of neural progenitor cell dynamics (PMID: 34556581). We have added the description regarding the overlapping and distinct functions of d-cysteine and d-serine to the Discussion (lines 418-436).

      (10) In the abstract, in lines 101 and 102, the authors mention "how d-serine contributes to cellular metabolism beyond neurotransmission remains largely unknown". In 2023, a paper in Stem Cell Reports by Roychaudhuri et al (PMID:37352848) showed that d and l-serine availability impacts lipid metabolism in the subventricular zone in mice, affecting proliferative properties of stem-cell derived neurons using a comprehensive lipidomics approach. There is no mention of this work even in the discussion section, as it bears directly on l and d-serine availability in neurons, which the authors are investigating. In the discussion section in lines 410-411, the authors mention the role of d-serine in neurogenesis, but surprisingly don't refer to the above reference. The role of d-serine in neurogenesis has been demonstrated in the Sultan et al (lines 855-857) and Roychaudhuri et al references.

      We thank the reviewer for highlighting these relevant studies. We have revised the statement in line 99 to avoid overgeneralization and to reflect that the metabolic roles of d-serine are incompletely understood rather than largely unknown. In addition, we have incorporated discussion of previous works, including the work by Roychaudhuri et al. (2023), into the Discussion section (lines 418-436) of our revised manuscript.

      (11) Both D-serine and the structurally similar stereoisomer D-cysteine (sulfur versus oxygen atom) have a bearing on 1C metabolism and the folate cycle. With reference to the folate cycle, Roychaudhuri et al in 2024 (PMID:39368613) have shown in rescue experiments in mice that supplementing a higher methionine diet provides folate cycle precursors to rescue the high insulin phenotype in SR-deficient mice. Since 1C metabolism is being discussed in this manuscript, the authors seem to overlook prior work in the field and not include it in their discussion, even when it is the same enzyme (SR) that synthesizes both serine and cysteine. Since the field of D-amino acid research is in its infancy, the authors must make it a point to include prior work related to D-serine at least, and not claim that it is not known. The known D-stereoisomers are not many, hence any progress in the area must include at least a discussion of the other structurally related stereoisomers.

      We are grateful to the reviewer for drawing our attention to this relevant study.

      We have now incorporated the findings from Roychaudhuri et al. 2024 into the Discussion. In that study, Srr-/- mice exhibited reduced levels of DNMT1 and DNMT3A, resulting in reduced DNA methylation activity. Notably, supplementation with a methyl-donor diet (containing choline, betaine, and methionine) restored the aberrant insulin phenotype in Srr-/- mice. These findings are relevant to our study, as DNA methylation depends on S-adenosyl-methionine (SAM), which is generated through one-carbon metabolism.

      We have expanded the Discussion to include the relationship between d-serine and d-cysteine, both of which are synthesized by Srr in the revised manuscript (lines 418-436).

      (12) Racemases (serine and aspartate) in general are promiscuous enzymes and known to synthesize other stereoisomers in addition to D-serine, D-cysteine, and D-aspartate. A few controls, like D-aspartate, D-cysteine, or even D-alanine must be included in their study to demonstrate the specific actions of D-serine, especially in the N2a cell treatment experiments. Cysteine and Serine are almost identical in structure (sulfur versus oxygen atom), and both are synthesized by serine racemase (published). Cysteine has also been very recently shown to inhibit tumor growth and neural progenitor cell proliferation. (PMIDs: 40797101 and 34556581). How the authors' work relates to the existing findings must be discussed, and this would put things in perspective for the reader.

      We thank the reviewer for this comment regarding the need to demonstrate the specificity of d-serine. As shown in Fig. S7A, d-serine, but not other d-amino acids commonly detected in mammals (Gonda et al., 2023), including d-aspartate, d-alanine, and d-proline, induced the cleavage of caspase-3 under the same experimental conditions, supporting the specific effect of d-serine in our system. We agree that d-cysteine shares structural similarity with d-serine and is also synthesized by Srr, suggesting potential functional overlap with d-serine. To place our findings in this context, we have added a paragraph in the Discussion (lines 418-436) describing the similarities of d-serine and d-cysteine. While both molecules may exert anti-proliferative effects, their underlying mechanisms appear to differ. Notably, supplementation with SAM, methionine, or glutathione did not rescue d-serine-induced growth inhibition in our system (Fig. S5J), suggesting that its effects are not primarily mediated through methylation or sulfur metabolic pathways. Instead, d-serine suppresses cellular proliferation by limiting mitochondrial l-serine availability and one-carbon metabolism. These observations highlight mechanistic divergence between d-serine and d-cysteine.

      Reviewer #2 (Public review):

      Summary:

      This study by Suzuki et al. reports an interesting stereo-selective role of D-serine in regulating one-carbon metabolism during neurodevelopment to adapt the functional transition, probably through the competition with mitochondrial transport of L-serine. The authors provide a multi-layered set of evidence, including metabolomics, enzyme assays, mitochondrial transport competition, and functional assays in immature/neural progenitor cells, to build up a conceptual integration of D-serine as both a neurotransmitter and a metabolic regulator in the central neural system, which raises a broad potential interest to the neuroscience and metabolism communities.

      Strengths:

      This work provides a conceptual advance that D-serine not only serves as a traditional neurotransmitter in the central neural system but also critically contributes to metabolic regulation of neural cells. The authors performed solid metabolomic assays to validate the suppressive effect of D-serine on the one-carbon metabolic pathway, providing some evidence that D-serine competitively inhibits mitochondrial serine transport, but not directly impairs SHMT2 enzymatic activity. All these data indicate a critical role of D-serine synthesis during neural maturation and suggest a potential translational strategy for targeting serine metabolism in neural tumors.

      Weaknesses:

      (1) The detailed mechanism by which D-serine competes with L-serine for its mitochondrial transport is not investigated. For example, although the authors made some discussion, they did not provide direct genetic or biochemical evidence linking these effects to the specific transporters, such as SFXN1.

      We thank the reviewer for this important comment regarding the mitochondrial l-serine transport mechanism. To address this point, we performed additional experiments using N2a cells in which Sfxn1 was knocked down by siRNA. Under semi-permeabilized cell conditions, we newly examined the effect of d-serine on mitochondrial L-serine transport.

      Interestingly, even under conditions where Sfxn1 expression was markedly suppressed, d-serine still inhibited mitochondrial d-serine transport. Given that SFXN1 is known to function redundantly with its paralogs (SFXN2–SFXN5) in mitochondrial serine transport (Kory et al., 2018), these findings suggest that d-serine may interfere with l-serine transport not only through SFXN1 but potentially through multiple members of the SFXN transporter family. These new data have been added as Fig. S3, and the corresponding results and discussion have been incorporated into the revised manuscript (lines 158–165 and 379-385).

      (2) Unlike tumor cells, where SHMT2 usually plays a predominant role in catalyzing serine/THF-derived one-carbon metabolism, normal cells may employ both SHMT1 and SHMT2 to do the work. Even under certain conditions that SHMT2-mediated one-carbon metabolism is suppressed, the activity of SHMT1 could be elevated for compensation. Thus, it is important to investigate whether D-serine affects SHMT1 activity or changes the balance between SHMT1- and SHMT2-mediated one-carbon metabolism. To this aim, the authors are strongly encouraged to perform a metabolic flux assay (MFA) by using 13C-labeled L-serine in the model cells in the presence and absence of D-serine.

      We thank the reviewer for this thoughtful comment regarding the potential contribution of cytosolic SHMT1 to one-carbon metabolism. As the reviewer notes, while mitochondrial SHMT2 is generally considered the predominant enzyme supporting one-carbon metabolism in proliferating or tumor cells, SHMT1 in the cytosol may function in a complementary manner in normal cells. In our experiments using primary cortical neurons, we indeed observed that the sensitivity to d-serine differed between immature and mature neuronal states (Fig. 4E). This observation suggests that the relative contribution of SHMT1- and SHMT2-mediated one-carbon metabolism may vary depending on the differentiation status of the cells.

      However, the primary focus of the present study was to investigate the mechanism underlying the anti-proliferative effect of d-serine in proliferative or undifferentiated neural cells. Our data demonstrate that d-serine inhibits one-carbon metabolism primarily by limiting mitochondrial l-serine availability through inhibition of mitochondrial l-serine transport. Therefore, a detailed analysis of the compensatory balance between SHMT1 and SHMT2 after disruption of mitochondrial serine transport falls beyond the central scope of the present study. Importantly, previous work by Miyamoto et al. (FEBS Journal, 2024) demonstrated that both SHMT1 and SHMT2 exhibit strong stereoselectivity for l-serine in their hydroxymethyl-transferase activity and do not utilize d-serine as a substrate. These findings make a direct effect of d-serine on hydroxymethyl-transferase activity of SHMT1 unlikely.

      Nevertheless, we agree that differences in the anti-proliferative effects of d-serine across cell types or differentiation states could reflect variations in cellular dependence on one-carbon metabolism or potential compensation by SHMT1. To address this point, we have expanded the Discussion section to clarify the possible contribution of SHMT1 and the limitations of the present study (lines 367–376). While isotope tracing analysis would be valuable for further dissecting compartmentalized one-carbon fluxes, such analyses would primarily address the relative contributions of SHMT1 and SHMT2 rather than the mitochondrial serine transport step that constitutes the central mechanism identified in this study.

      (3) A defect in serine-derived one-carbon metabolism may cause multiple cellular stress responses. It is valuable to detect whether cellular NADPH/NADH, GSH, or ROS is altered before and after D-serine treatment.

      We appreciate this insightful comment regarding potential cellular stress responses associated with impaired one-carbon metabolism. Consistent with the reviewer’s suggestion, we examined markers related to redox status. Transcriptomic analysis revealed compensatory changes in genes involved in mitochondrial metabolic function, including components of the NADH dehydrogenase complex (Fig. 2IJ and Fig. S4A), suggesting metabolic adaptation to d-serine treatment. In response to the reviewer’s comment, we measured GSH levels in NPCs and found that d-serine treatment led to a reduction of GSH (new Fig. S7F), indicating altered redox balance. However, supplementation with exogenous GSH did not rescue d-serine-induced cell death (Fig. S7E). These results suggest that while d-serine induces changes in cellular redox status, including GSH depletion, redox imbalance alone is unlikely to be the primary driver of cell death in this context.

      (4) The physiological relevance between D-serine and neural cell maturation/death should be further tested and discussed, since the dosage of D-serine used in the in vitro assay is much higher than that in physiological conditions.

      We thank the reviewer for this comment regarding the physiological relevance of the d-serine concentrations used in our study. We agree that the concentrations of d-serine required to compete with l-serine for mitochondrial transport are higher than those typically observed under physiological conditions, which we acknowledged in the Discussion of our manuscript (lines 386-388). Importantly, however, in vivo, d-serine levels are tightly regulated in a spatiotemporal manner during brain development, with low levels during embryonic and early postnatal stages and increased levels upon neural maturation. This temporal regulation coincides with the transition from proliferative neural progenitor states to differentiated neurons, thereby limiting the potential for d-serine to interfere with one-carbon metabolism during periods of active cell proliferation. Thus, while the concentrations used in vitro may exceed physiological levels, they allow us to uncover a latent metabolic effect of d-serine that may become relevant under specific cellular or developmental contexts. To clarify these points, we have revised the Discussion (lines 388-395) in the revised manuscript to more explicitly address the relationship between d-serine dosage and physiological relevance.

      Reviewer #3 (Public review):

      Summary:

      This manuscript presents a comprehensive and well-executed investigation into the metabolic role of D-serine in the central nervous system. The authors provide solid evidence that D-serine competitively inhibits mitochondrial L-serine transport, thereby impairing one-carbon metabolism. This stereoselective mechanism reduces glycine and formate production, suppresses cellular proliferation, and induces apoptosis in immature neural cells and glioblastoma stem cells. Developmental analyses further reveal a physiological enantiomeric shift in serine metabolism during neurogenesis, aligning with the transition from proliferation to maturation. Overall, the study bridges developmental neurobiology, cancer metabolism, and amino acid transport, uncovering a previously unrecognized metabolic function of D-serine beyond its role in neurotransmission.

      Strengths:

      (1) The discovery that D-serine inhibits one-carbon metabolism by competing for mitochondrial L-serine transport-rather than through enzymatic inhibition or receptor-mediated signaling-represents a significant and previously underappreciated mechanism. This finding has broad implications for understanding metabolic regulation during neurodevelopment and offers potential relevance for targeting metabolic vulnerabilities in cancer.

      (2) The authors integrate metabolomics, mitochondrial transport assays, molecular dynamics simulations, genetic and pharmacologic perturbations, transcriptomics, and both in vitro and ex vivo models. The breadth of experimental approaches, combined with the coherence of the findings across systems, provides strong support for the central conclusions and enhances the overall impact of the study.

      (3) The temporal shift in D-/L-serine levels during neurodevelopment is elegantly linked to the transition from proliferative to mature neuronal states. The selective vulnerability of neural progenitors and tumor cells-contrasted with the resistance of mature neurons-highlights a biologically meaningful and potentially targetable metabolic distinction.

      Weaknesses:

      (1) While the authors attribute D-serine's metabolic effects to competition with mitochondrial L-serine transport, the specific identity of the transporter(s) mediating this process remains undefined. This represents a meaningful mechanistic gap, as the central conclusion depends on D-serine limiting mitochondrial L-serine availability to inhibit one-carbon metabolism.

      We thank the reviewer for this insightful comment regarding the identity of the mitochondrial l-serine transporter. As this concern overlaps with the point raised by Reviewer 2 (Weakness 1), we refer the reviewer to our response there for a detailed description of the additional experiments performed. Briefly, we conducted new experiments using siRNA-mediated knockdown of Sfxn1 in N2a cells and examined mitochondrial l-serine transport under semi-permeabilized conditions. Notably, even with marked suppression of Sfxn1 expression, d-serine continued to inhibit mitochondrial l-serine transport. Given that SFXN family members (SFXN1–SFXN5) are reported to function redundantly in mitochondrial serine transport (Kory et al., 2018), these findings suggest that d-serine may interfere with l-serine transport not only via SFXN1 but potentially across multiple SFXN paralogs. These results have been incorporated into the revised manuscript and are presented in Fig. S3, with the corresponding discussion added to the Results and Discussion sections (lines 158–164 and 379-385).

      (2) The effective concentrations of D-serine used in vitro (IC<sub>50</sub> ≈ 1-2 mM) exceed typical brain levels (~0.3 mM). While the authors acknowledge this, a more focused discussion on whether higher local D-serine concentrations could arise in specific microenvironments - such as synaptic compartments, tumor niches, or pathological states-would help contextualize the in vitro findings and strengthen their physiological relevance. For example, disruptions in D-serine clearance or altered expression of serine racemase and transporters in disease contexts could lead to localized accumulation. Moreover, differences between extracellular and intracellular D-serine pools - and the mechanisms governing their regulation - may further influence its metabolic impact in vivo.

      We appreciate this insightful comment regarding the physiological relevance of the d-serine concentrations used in vitro. We agree that the effective concentrations observed in our assays exceed typical bulk brain levels. To address this point, we have expanded the Discussion (lines 386-399) to consider conditions under which locally elevated d-serine concentrations may arise in vivo. These additions provide a more nuanced interpretation of the relationship between the concentrations used in vitro and the potential physiological contexts in which d-serine may exert metabolic effects.

      (3) While the manuscript focuses on neural stem/progenitor cells and neural tumors, it remains unclear whether the anti-proliferative effects of D-serine are specific to neural lineages or extend to other highly proliferative non-neural cell types. A brief discussion addressing this point would help clarify the scope of D-serine's metabolic impact and whether its mechanism of action reflects a unique vulnerability in neural cells or a more general feature of proliferative metabolism. This distinction is particularly relevant for assessing the broader therapeutic potential of targeting mitochondrial L-serine transport.

      We thank the reviewer for this comment regarding the potential generality of the anti-proliferative effects of d-serine. We agree that it is important to clarify whether the observed effects are specific to neural lineages or reflect a broader vulnerability of proliferative cells. In our study, we focused on neural progenitor cells and neural tumour models. However, the mechanism identified here, namely limitation of mitochondrial L-serine availability, targets a fundamental metabolic pathway that supports cell proliferation. Therefore, it is possible that similar effects may extend to other highly proliferative cell types beyond the neural lineage. To address this point, we have expanded the Discussion (lines 408-410) to clarify that the observed effects of d-serine may reflect a general metabolic vulnerability associated with proliferative states, while also noting that the degree of sensitivity is likely to depend on cell-type-specific reliance on mitochondrial one-carbon metabolism.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor issues:

      The authors mention terms like neural tissue (line 109) and neural tumor cells (line 182). Neural tissue can mean anything under the sun. They need to mention the specific tissue being studied or investigated.

      We have changed the terms.

      Reviewer #2 (Recommendations for the authors):

      (1) The figure items were not well organized, and the legends were difficult to read since they apparently lacked key information. For example, in Figure 2D, did the author intend to show SHMT2 activity? It is not clear how this assay was performed (in vitro or in vivo experiment?). Additionally, more detailed information should be provided in the Methods section.

      We have improved the organization of figures and revised figure legends for better readability.

      (2) The writing of the manuscript should be significantly improved, using a professional editing service, in order to increase the readability.

      We appreciate the reviewer’s comment. The manuscript was professionally edited prior to submission. Nevertheless, we have made targeted revisions to the Introduction and Discussion sections to improve overall readability. We hope that these revisions have improved the clarity of the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) The core mechanism centers on competition for mitochondrial L-serine transport, yet the identity of the transporter(s) involved remains speculative. While Kory et al. (2018) identified SFXN1 as a mitochondrial L-serine transporter, this connection is not directly addressed in the current study. It would strengthen the manuscript to clarify whether SFXN1 or related isoforms are expressed in the neural cell models used and whether their expression patterns correspond with the observed D-serine sensitivity. Even if functional validation is beyond the current scope, a more detailed discussion of potential transporter candidates and the limitations of existing data would provide important mechanistic context and help frame future directions.

      We thank the reviewer for this insightful comment regarding the mitochondrial l-serine transporter and the potential involvement of SFXN family members. As this concern overlaps with the point raised by Reviewer 2 (Weakness 1), we refer the reviewer to our response there for a detailed description of the additional experiments performed.

      Briefly, we performed siRNA-mediated knockdown of Sfxn1 in N2a cells and examined mitochondrial l-serine transport under semi-permeabilized conditions. Even under conditions of marked suppression of Sfxn1 expression, d-serine continued to inhibit mitochondrial l-serine transport. Given that members of the SFXN family have been reported to function redundantly in mitochondrial serine transport (e.g., Kory et al., 2018), these findings suggest that d-serine may affect l-serine transport not only through SFXN1 but potentially across multiple SFXN paralogs.

      These new results have been incorporated into the revised manuscript (Fig. S3), and the relevant discussion has been expanded to clarify the potential roles of SFXN family transporters and the current limitations in defining the exact transporter responsible (lines 158–165 and 379-385).

      (2) The authors use ex vivo brain slice cultures with tumor xenografts to demonstrate tissue-level relevance, which is a valuable strength of the study. However, additional context would enhance its translational significance. It would be helpful to discuss whether in vivo D-serine administration (e.g., ICV or systemic) is feasible and safe, especially given the high concentrations required in vitro. Briefly addressing whether genetic models, such as Srr knockout mice, support a role for D-serine in tumor progression or neurodevelopment would also strengthen the interpretation.

      We thank the reviewer for this important comment regarding the translational relevance of d-serine administration in vivo. To address this point, we performed additional exploratory experiments using a subcutaneous tumour xenograft model in nude mice. Because the inhibitory effect of d-serine on one-carbon metabolism becomes evident under l-serine–limited conditions, tumor-bearing mice were fed an l-serine/glycine–deficient diet and administered d-serine in drinking water.

      However, when d-serine was provided at high concentrations (≥ 500 mM) in drinking water, the mice exhibited marked behavioral abnormalities, including increased aggression and other abnormal behaviors. Due to these adverse effects, the experiment was ethically terminated. While our study focuses on the inhibitory effect of d-serine on one-carbon metabolism, d-serine is also a physiological co-agonist of the NMDA receptors, and therefore high systemic concentrations may influence neuronal excitability in vivo. These observations suggest that the concentrations required to achieve antitumor effects may be associated with significant neurological side effects.

      Based on these findings, we consider that direct administration of d-serine itself may have limited therapeutic applicability as an antitumor reagent. Instead, future development of derivatives or strategies that retain the metabolic inhibitory effect while minimizing NMDA receptor–mediated effects may be required.

      Regarding serine racemase (Srr) knockout models, xenograft tumour experiments would require an immunodeficient background, which makes the generation and use of such compound models technically challenging. Therefore, we did not pursue this approach in the present study. We have incorporated these considerations into the revised Discussion (lines 413-417) to clarify the translational implications and current limitations of our findings.

      Bella DJD, Habibi E, Stickels RR, Scalia G, Brown J, Yadollahpour P, Yang SM, Abbate C, Biancalani T, Macosko EZ, Chen F, Regev A, Arlotta P. 2021. Molecular logic of cellular diversification in the mouse cerebral cortex. Nature 595:554–559. DOI: https://doi.org/10.1038/s41586-021-03670-5, PMID: 34163074

      Bernal A, Arranz L. 2018. Nestin-expressing progenitor cells: function, identity and therapeutic implications. Cellular and Molecular Life Sciences 75:2177–2195. DOI: https://doi.org/10.1007/s00018-018-2794-z, PMID: 29541793

      Bott CJ, Johnson CG, Yap CC, Dwyer ND, Litwa KA, Winckler B. 2019. Nestin in immature embryonic neurons affects axon growth cone morphology and Semaphorin3a sensitivity. Molecular Biology of the Cell 30:1214–1229. DOI: https://doi.org/10.1091/mbc.e18-06-0361, PMID: 30840538

      Graham V, Khudyakov J, Ellis P, Pevny L. 2003. SOX2 Functions to Maintain Neural Progenitor Identity. Neuron 39:749–765. DOI: https://doi.org/10.1016/s0896-6273(03)00497-5, PMID: 12948443

      Lendahl U, Zimmerman LB, McKay RDG. 1990. CNS stem cells express a new class of intermediate filament protein. Cell 60:585–595. DOI: https://doi.org/10.1016/0092-8674(90)90662-x, PMID: 1689217

      Thakurela S, Tiwari N, Schick S, Garding A, Ivanek R, Berninger B, Tiwari VK. 2016. Mapping gene regulatory circuitry of Pax6 during neurogenesis. Cell Discovery 2:15045. DOI: https://doi.org/10.1038/celldisc.2015.45, PMID: 27462442

    1. eLife Assessment

      This useful study investigates how plasticity and homeostatic adaptation can lead to the emergence of synchronization patterns associated with different sleep phases. The ideas are novel and interesting, but the evidence at present remains incomplete. Further work is needed to determine whether the reported states persist for larger system sizes, longer integration times, and how robust the results are to different initial conditions.

    2. Reviewer #1 (Public review):

      Summary:

      The authors aim to understand how changes in the balance between excitatory and inhibitory interactions influence the stability and reorganization of network connections. To address this question, they extend a coupled-phase-oscillator model by adding plasticity rules. The central finding is that stronger inhibitory interactions lead to relatively stable and desynchronized network dynamics, whereas weaker inhibitory interactions produce a bistable regime in which intermediate-strength connections fluctuate while stronger connections are preserved.

      Strengths:

      This study offers a simple theoretical framework for linking network state, coupling stability, and reorganization. The model produces clear qualitative results, showing that different dynamical regimes are associated with different balances of excitatory and inhibitory interactions. This could be useful as a conceptual starting point for considering how network states may regulate the stability and flexibility of connections. The manuscript also explores several model parameters.

      Weaknesses:

      The evidence is incomplete in supporting the biological interpretations. The model is a highly simplified coupled-phase-oscillator system and does not directly represent spiking activity, membrane potentials, synaptic currents, conduction delays, cellular excitability, or detailed biological plasticity mechanisms. Although the authors clarify that the model units are not actual neurons or synapses, the discussion often interprets the results in terms of neuronal inhibition, synaptic stability, sleep-related reorganization, and preservation of strong biological connections. This creates a gap between the abstract model and the biological conclusions. In particular, the manuscript does not sufficiently discuss what biological oscillatory activity the modeled phases are intended to represent, such as population-level activity reflected in electroencephalography or local field potentials. In several places, the manuscript appears to assume that neurons can generally be treated as oscillators, but this is not always a valid assumption. The authors should more clearly distinguish between rhythmic or phase-like activity at the population level and the dynamics of individual neurons, and should frame the model more cautiously as a phenomenological description of collective synchronization rather than a mechanistic model of spiking neuronal circuits.

      There are also important methodological limitations. Although the manuscript presents the model equations, parameter values, time step, simulation duration, and coupling update rules, several other essential details are not clearly specified, including the number of simulation runs, the procedure for setting initial conditions, and the numerical method used to solve the ordinary differential equations. Critically, technical details such as the integration scheme, solver settings, initialization procedure, and random seed handling are essential for reproducibility. Because the main findings depend on the interaction between phase dynamics and adaptive coupling, even small implementation differences could affect the reported dynamical regimes and coupling fluctuations.

      A further concern is the presentation of the mathematical formulation. Several equations appear to contain notation errors and inconsistencies, making it difficult to follow the exact model definition. The authors should carefully revise the mathematical notation throughout the manuscript to ensure that the model can be understood and reproduced unambiguously.

      Overall, the study provides a useful but limited theoretical account of how network dynamics may regulate coupling stability and reorganization. The results support the internal behavior of the proposed model, but the broader biological claims are not yet fully convincing. The likely impact of the work is therefore mainly conceptual: it may stimulate further modeling studies, but additional methodological detail, stronger justification of the modeling assumptions, and comparison with more biologically grounded models would be needed before the conclusions can be applied confidently to neuronal circuit dynamics or sleep-related synaptic reorganization.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the impact of plasticity mechanisms in an excitation-inhibition (EI) network model on the emergence of synchronization patterns that the authors associate with different sleep phases. The model consists of an EI Kuramoto network in which recurrent excitatory couplings evolve according to Hebbian and homeostatic adaptation rules. Through numerical simulations, the authors analyze how these plasticity mechanisms modify both the collective dynamics and the structure of the coupling matrix.

      Strengths:

      The topic addressed in the manuscript is timely and potentially relevant, as understanding the interplay between synaptic adaptation and collective neural dynamics remains an important challenge in theoretical neuroscience.

      Weaknesses:

      In its current form, the work suffers from substantial conceptual, methodological, and technical limitations that significantly weaken the conclusions.

      From a biological perspective, the model is highly abstract and qualitative. The connection between the model variables and the physiological processes that the authors aim to describe remains unclear. Consequently, the manuscript does not provide sufficient evidence to support biologically meaningful conclusions regarding sleep dynamics. In my opinion, the work is more naturally positioned within the framework of theoretical or computational dynamical systems than within the scope of a biology-oriented journal.

      From a mathematical and dynamical-systems perspective, the analysis is incomplete, and several important technical aspects are either missing or inadequately addressed. In particular, the characterization of the dynamical regimes is often imprecise, the numerical evidence is not sufficiently robust, and little effort is made to interpret the results within the broader context of synchronization theory, adaptive networks, or collective dynamics.

      More specifically:

      (1) The biological interpretation of the model variables is ambiguous throughout the manuscript. At several points, the authors suggest that individual oscillators should not be interpreted as neurons but rather as abstract biological units (lines 78-82, 86-89, 418-420). However, other parts of the manuscript refer to the coupling matrix entries, particularly $J_{ee}$, as synaptic weights (e.g., line 118). These two interpretations are not obviously compatible. If the oscillators represent coarse-grained or abstract units, the biological meaning of the adaptive couplings should be carefully justified. More generally, the manuscript lacks a clear discussion of what aspects of neural circuits are captured by the model and which aspects are intentionally neglected.

      (2) More fundamentally, the manuscript inherits the well-known limitations associated with interpreting Kuramoto oscillators as neural elements. Kuramoto phase oscillators provide a minimal description of synchronization phenomena, but they do not explicitly represent membrane dynamics, firing rates, spiking activity, synaptic currents, or realistic neuronal timescales.

      Under certain assumptions, Kuramoto-like models can be rigorously derived from more detailed neuronal models through phase-reduction techniques (see, for instance, Chapter 10 of Izhikevich's \textit{Dynamical Systems in Neuroscience}). However, the authors do not employ such a reduction procedure, nor do they establish a formal connection between their model variables and the dynamics of neuronal populations. As a consequence, the biological interpretation of the model remains unclear.

      The authors should therefore explicitly discuss these limitations and carefully justify why the synchronization patterns observed in such a highly reduced model can be related to neural sleep states. At present, the biological interpretation appears considerably stronger than what the model itself can support, for instance, the claims in lines 321-322 or 330-337. In particular, it remains unclear whether the reported dynamical regimes should be interpreted as genuine mechanisms underlying sleep rhythms or merely as generic synchronization phenomena arising in adaptive oscillator networks.

      (3) The numerical methodology raises serious concerns regarding the robustness of the reported results. According to the Methods section, simulations are performed using only $N=100$ oscillators and integration times of approximately 50 time units. Such choices may be sufficient for illustrative purposes but are generally inadequate for drawing conclusions about asymptotic collective behavior in adaptive dynamical systems. Finite-size fluctuations can strongly affect synchronization measures, and adaptive networks are well known to exhibit extremely long transients, metastability, and slow convergence processes. No systematic finite-size analysis is provided, nor is there any demonstration that the reported states persist for larger system sizes or longer integration times.

      (4) The use of the term "bistable regime" to describe the dynamics shown in Figure 1C is incorrect. A bistable regime refers to a parameter region in which multiple attractors coexist and the asymptotic state depends on the initial condition. The figure instead presents a single trajectory displaying oscillatory dynamics. No evidence is provided for the coexistence of attractors, nor are multiple initial conditions explored. Furthermore, the displayed time series are too short to determine whether the observed dynamics correspond to a stable limit cycle, quasiperiodic motion, intermittent behavior, or a long transient approaching another attractor. The authors should perform a proper dynamical characterization of this regime. Similar collective oscillatory states have been extensively studied in synchronization and adaptive-network models and should be discussed in relation to the existing literature.

      In particular, several time series shown in the manuscript (e.g., Figures 1C and 2F) exhibit trends that suggest the possibility of unresolved transient dynamics. The authors should demonstrate convergence of the reported regimes by substantially extending simulation times and by performing finite-size analyses. Simulations with at least one order of magnitude more units ($N\gtrsim 1000$) and integration times sufficient to establish asymptotic behavior would be expected in a study whose main claims rely on collective dynamical phenomena.

      (5) The manuscript lacks several standard tools routinely employed in the analysis of nonlinear dynamical systems. The conclusions are largely based on visual inspection of time series and order parameters. However, no bifurcation analysis, stability analysis, phase-space characterization, attractor reconstruction, or systematic exploration of parameter dependence is provided. As a consequence, many of the identified "phases" or "regimes" remain only qualitatively described. A more rigorous dynamical-systems treatment would substantially strengthen the work and would help distinguish genuine asymptotic states from finite-size or transient phenomena.

      (6) A substantial fraction of the results appears to extend the authors' previous work by incorporating plastic adaptation mechanisms. While incremental advances are acceptable, the manuscript would benefit from a broader theoretical context. The discussion is heavily centered on previous studies by the same authors, whereas there exists an extensive literature on synchronization, adaptive networks, neural mass models, balanced EI systems, and sleep-related oscillations that is largely absent from the discussion. The novelty and significance of the present contribution would be easier to assess if the results were more carefully compared with alternative theoretical approaches.

    1. eLife Assessment

      This important study introduces a non-perturbative pulse-labeling strategy for yeast nuclear pore complexes (NPCs), employing a nanobody-based approach in order to selectively capture Nup84-containing complexes for imaging and biochemical analysis. The data convincingly demonstrate that a short induction period (20 minutes to 1 hour) yields a strong and sustained signal, enabling affinity purification that faithfully recapitulates the endogenous Nup84 interactome. This tool offers a powerful framework for investigating NPC dynamics and associated interactomes through both imaging and biochemical assays.

    2. Reviewer #1 (Public review):

      Summary:

      The authors present a nanobody-based pulse-labeling system to track yeast NPCs. Transient expression of a nanobody targeting Nup84 (fused to NeonGreen or an affinity tag) permits selective visualization and biochemical capture of NPCs. Short induction effectively labels NPCs, and the resulting purifications match those from conventional Nup84 tagging. Crucially, when induction is repressed, dilution of the labeled pool through successive cell cycles allows the visualization of "old" NPCs (and potentially individual NPCs) providing a powerful view of NPC lifespan and turnover without permanently modifying a core scaffold protein.

      Strengths:

      (1) A brief expression pulse labels NPCs, and subsequent repression allows dilution-based tracking of older (and possibly single) NPCs over multiple cell cycles.

      (2) The affinity-purified complexes closely match known Nup84-associated proteins, indicating specificity and supporting utility for proteomics.

      Weakness:

      Reliance on GAL induction introduces metabolic shifts (raffinose → galactose → glucose) that could subtly alter cell physiology or the kinetics of NPC assembly. As acknowledged by the authors, alternative induction systems (e.g., β-estradiol-responsive GAL4-ER-VP16) could be implemented as a way to avoid carbon-source changes.

      Comments on revised version.

      The authors have thoughtfully addressed all of my concerns. In particular, they have updated the proteomic analysis in Figure 1I, showing that they recover most NPC components (including basket Nups), including non-NPC proteins as controls, and providing all data as a supplementary table. These changes strengthen the authors conclusion and improve transparency. I have no further recommendations and congratulate the authors for their exciting work.

    3. Reviewer #2 (Public review):

      Summary:

      This preprint describes a practical and useful approach for labeling and tracking NPCs in situ, using a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity. Useful applications including timelapse imaging, affinity purification, and proximity labeling are envisioned.

      Strengths:

      Clever use of a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity.

    4. Reviewer #3 (Public review):

      Summary:

      Submitted to the Tools and Resources series, this study reports on the use of a single-domain antibody targeting the nucleoporin Nup84 to probe and track NPCs in budding yeast. The authors demonstrate their ability to rapidly label or pull down NPCs by inducing the expression of a tagged version of the nanobody (Fig. 1).

      Strengths:

      This tool's main strength is its versatility as an inexpensive, easy-to-set-up alternative to metabolic labelling or optical switching. This same rationale could, in principle, be applied to the study of other multiprotein complexes using similar strategies, provided that single-chain antibodies are available.

      Weaknesses:

      This approach has no inherent weaknesses, but it would be useful to verify in the future that this pulse labelling strategy can also be used to detect assembly intermediates, structural variants, or damaged NPCs, e.g. NPC clusters formed in some nucleoporin mutants.

      Overall, the data clearly shows that Nup84 nanobodies are a valuable tool for imaging NPC dynamics and investigating their interactomes through affinity purification.

      Comments on revised version.

      None at this stage.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors present a nanobody-based pulse-labeling system to track yeast NPCs. Transient expression of a nanobody targeting Nup84 (fused to NeonGreen or an affinity tag) permits selective visualization and biochemical capture of NPCs. Short induction effectively labels NPCs, and the resulting purifications match those from conventional Nup84 tagging. Crucially, when induction is repressed, dilution of the labeled pool through successive cell cycles allows the visualization of "old" NPCs (and potentially individual NPCs), providing a powerful view of NPC lifespan and turnover without permanently modifying a core scaffold protein.

      Strengths:

      (1) A brief expression pulse labels NPCs, and subsequent repression allows dilution-based tracking of older (and possibly single) NPCs over multiple cell cycles.

      (2) The affinity-purified complexes closely match known Nup84-associated proteins, indicating specificity and supporting utility for proteomics.

      We thank the reviewer for this evaluation

      Weaknesses:

      (1) Reliance on GAL induction introduces metabolic shifts (raffinose -> galactose -> glucose) that could subtly alter cell physiology or the kinetics of NPC assembly. Alternative induction systems (e.g., β-estradiol-responsive GAL4-ER-VP16) could be discussed as a way to avoid carbon-source changes.

      Indeed, this could be an improvement, and we mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3.

      (2) While proteomics is solid, a comprehensive supplementary table listing all identified proteins (with enrichment and statistics) would enhance transparency.

      Indeed, we now provide source data showing LFQ intensities, fold-enrichment and statistics for all detected proteins.

      (3) Importantly, the authors note that the method is particularly useful "in conditions where direct tagging of Nup84 interferes with its function, while sub-stoichiometric nanobody binding does not." After this sentence, it would be valuable to add concrete examples, such as experiments examining NPC integrity in aging or stress conditions where epitope tags can exacerbate phenotypes. These examples will help readers identify situations in which this approach offers clear advantages.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants, and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks, and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #2 (Public review):

      Summary:

      This preprint describes a practical and useful approach for labeling and tracking NPCs in situ. While useful applications including timelapse imaging, affinity purification, or proximity labeling are envisioned, addressing some outstanding technical questions would give a clearer picture of the sensitivity and temporal resolution of this approach.

      Strengths:

      Clever use of a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity.

      We thank the reviewer for this evaluation

      Weaknesses:

      The decrease in nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over during this time. However, it is also possible that the nanobody: Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #3 (Public review):

      Summary:

      Submitted to the Tools and Resources series, this study reports on the use of a single-domain antibody targeting the nucleoporin Nup84 to probe and track NPCs in budding yeast. The authors demonstrate their ability to rapidly label or pull down NPCs by inducing the expression of a tagged version of the nanobody (Figure 1).

      Strengths:

      This tool's main strength is its versatility as an inexpensive, easy-to-set-up alternative to metabolic labelling or optical switching. This same rationale could, in principle, be applied to the study of other multiprotein complexes using similar strategies, provided that single-chain antibodies are available.

      We thank the reviewer for this evaluation

      Weaknesses:

      This approach has no inherent weaknesses, but it would be useful for the authors to verify that their pulse labelling strategy can also be used to detect assembly intermediates, structural variants, or damaged NPCs.

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure but consider such studies to be beyond the scope of this study.

      Overall, the data clearly show that Nup84 nanobodies are a valuable tool for imaging NPC dynamics and investigating their interactomes through affinity purification.

      We thank the reviewer again for the constructive feedback and thoughts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 1A, and although it is partially mentioned in the legend, it would be helpful to indicate precisely when cells are grown in raffinose, when galactose is added for induction, and when glucose is used to terminate expression.

      We included “galactose” and “glucose” to Panel A to indicate induction and termination of expression, respectively.

      (2) Related to the previous point, consider mentioning the GAL4-ER-VP16 (ADGEV) estradiol-inducible system as an optional strategy to avoid carbon shifts and potentially reduce cell-to-cell variability.

      We mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3

      (3) Add a brief sentence explaining that the ZZ tag is derived from Protein A and binds IgG Fc.

      This information is now added on p.2

      (4) The statement "all Nups significantly coenriched with VHH[Nup84]-ZZ..." is likely inaccurate, since not all Nups are labeled in panels F-H, and some basket components are missing in panel I (particularly basket components such as Nup60, Nup1). Consider revising to "most Nups significantly coenriched...". In panel I, please include a clearly non-enriched protein as a visual reference for the color scale.

      We are very grateful to the reviewer for pointing this out. We accidentally used a faulty filtering on the dataset to generate figure panel I, omitting several Nups that were reproducibly found in all replicas. All Nups, except for Gle1 and Pom33, were detected reproducibly.

      We have made the following adjustments to the figure panel and accompanying text:

      In Fig. 1I, we included the missing Nups and 5 proteins that co-purified with VHH[Nup84] but not specifically enriched, as the reviewer suggested. They cluster in a separate group and their abundance is not going up in time. We randomly selected these 5 proteins from the list of genes that were reproducibly found in all four timepoints.

      For clarity, we removed the NTRs

      We changed the text to “we found that all Nups, except Gle1 and Pom33, significantly coenriched with VHH[Nup84]-ZZ” on p.2.

      We updated the methods section, describing the clustering method and how we selected the 5 random proteins

      (5) Provide a supplementary spreadsheet with LFQ intensities, fold-enrichment, and statistics for all detected proteins. This will address questions about missing Nups and support transparency.

      This information is now added as Source data Figure 1.

      (6) Directly after the statement "Amongst others this is useful in conditions where direct tagging of Nup84 interferes with its function, while sub stoichiometric nanobody binding does not," it would be useful to include concrete instances, such as stress or aging conditions, where Nup84 tagging may sensitize NPC integrity.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      (7) In panels K and L, since individual points correspond to biological replicates, overlaying a box plot obscures much of the data. Consider overlaying the means per replicate instead of box plots: see the "SuperPlots" approach for a clear explanation of how to present this (PMID: 32346721).

      We thank the reviewer for the “SuperPlots” suggestion, and we agree that representing the data in this way improves the visualization of individual biological replicates. We have updated the summarizing overlay in figures in panel K and L to represent the means per replicate instead of boxplots.

      (8) I spotted a few typos ("Lasty" ? "Lastly"; "in maintained" vs. "is maintained").

      Thank you, these are corrected

      Overall, this is a neat, well-executed methodological advance with clear value to the NPC field and potentially other complex assemblies. I look forward to seeing a revised version.

      Thank you!

      Reviewer #2 (Recommendations for the authors):

      Based on the recent structural analyses and NPC modeling using this nanobody, how accessible is the Nup84 epitope expected to be within the fully assembled NPC? While the data shown indicate that nanobody labeling of NPCs is readily detectable, stating this clearly would help motivate the approach and interpret the resulting data.

      We now included such a statement in the introduction on p.1.

      The decrease of nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over due to cell division during this time window. However, it is also possible that nanobody:Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      Reviewer #3 (Recommendations for the authors):

      (1) As mentioned above, to assess the general relevance of this tool, it would be informative to verify whether the VHH[Nup84] nanobody can access and detect NPC species under conditions that challenge their structural organization or biogenesis, for example, in nucleoporin mutants or under stress. The authors could, for instance, analyze the localization of VHH[Nup84] in yeast strains harboring clustered NPCs (nup133Δ), or following stresses known to impact NPC organization (e.g., osmotic stress or energy depletion; PMID: 34762489).

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure and tried to include such data. Unfortunately, this was not successful, and further efforts are beyond the scope of his study. Following the reviewer’s suggestion, we expressed VHH[Nup84] in nup133∆N (nup133∆2-300) (Doye, Wepf, and Hurt 1994) following the experimental set-up in panel A and examined its localization. However, at t=2hrs hardly any nanobody signal was detectable in nup133∆N (see Author response image 1, upper panel A) and only after overnight expression nanobody-labelled NPC clusters are detectable (bottom panel A). Considering that expression levels of free mNG are also lower at t=2hrs in nup133∆N cells compared to WT cells (Author response image 1, panel B), it appears that protein expression under the Gal system is generally reduced in a nup133∆N background. These expression level differences between nup133∆N and WT preclude statements about the accessibility of the Nup84 epitope in nup133∆N. We note that nup133∆N cells do not have general mRNA export defects (Doye, Wepf, and Hurt 1994), so other inducible systems may be better suited for such analysis.

      Author response image 1.

      Expression level differences in WT and Nup133∆N cells. Left: localization of VHH[Nup84]-mNG in Nup133∆N cells at t=2hr following a 20-minute induction pulse and after overnight 0.5% galactose (ON) induction. Right: mNG levels in WT and Nup133∆N cells at t=2hr following a 20-minute induction pulse. Brightness/contrast settings are identical between the two panels. All panels are sum slices projections from 30 z-slices of 0.1µm. Scale bar = 5 µm.

      (2) Since outer rings are found on both sides of NPCs (i.e., the cytoplasmic and nuclear faces), could the authors indicate whether the VHH[Nup84] nanobody can enter the nucleus and probe the nuclear outer rings? Along these lines, it would be useful to provide a summary of the structural organization of NPCs in the introduction.

      Thank you, we have added a sentence on the localization of Nup84 in NPCs in the introduction. Based on what is known about influx (nuclear transport receptor-independent nuclear entry) of proteins with similar size and surface properties (Popken et al. 2015; Timney et al. 2016), the nanobody can rapidly enter the nucleus and hence bind Nup84 on both the nuclear and cytoplasmic side. We have no data to answer if binding might initially be biased towards cytosolic VHH[Nup84] binding the cytoplasmic outer rings.

      (3) The authors state that VHH[Nup84] and direct Nup84 detection are indistinguishable (p. 2). Could they provide images of the endogenously tagged Nup84-GFP strain for comparison?

      We have now included a pairwise comparison in a Figure 1 – supplement 1.

      Minor corrections:

      (1) There are a few typos that need correcting: 'Nup84Δ' (p. 1; should read 'nup84Δ') and 'promotor' (p. 2; should read 'promoter').

      Thank you, these are corrected

      (2) The reference 'Veldsink et al. 2025' (quoted in the PunctaFinder analysis description on page 8) does not appear in the References section.

      Thank you, these are corrected.

      We thank the reviewer again for the constructive feedback and thoughts.

      References

      Doye, V., R. Wepf, and E. C. Hurt. 1994. 'A novel nuclear pore protein Nup133p with distinct roles in poly(A)+ RNA transport and nuclear pore distribution', EMBO J, 13: 6062-75.

      Khmelinskii, Anton, Philipp J. Keller, Holger Lorenz, Elmar Schiebel, and Michael Knop. 2010. 'Segregation of yeast nuclear pores', Nature, 466: E1-E1.

      Popken, Petra, Ali Ghavami, Patrick R. Onck, Bert Poolman, and Liesbeth M. Veenhoff. 2015. 'Size-dependent leak of soluble and membrane proteins through the yeast nuclear pore complex', Molecular Biology of the Cell, 26: 1386-94.

      Timney, Benjamin L., Barak Raveh, Roxana Mironska, Jill M. Trivedi, Seung Joong Kim, Daniel Russel, Susan R. Wente, Andrej Sali, and Michael P. Rout. 2016. 'Simple rules for passive diffusion through the nuclear pore complex', Journal of Cell Biology, 215: 57-76.

      Zsok, J., F. Simon, G. Bayrak, L. Isaki, N. Kerff, Y. Kicheva, A. Wolstenholme, L. E. Weiss, and E. Dultz. 2024. 'Nuclear basket proteins regulate the distribution and mobility of nuclear pore complexes in budding yeast', Mol Biol Cell, 35: ar143.

    1. eLife Assessment

      This valuable study presents a comparative analysis of the transcriptomic features underlying C. elegans longevity, providing insights into how different changes in gene expression can promote longevity. The authors present solid evidence with analysis and selected functional validation showing that some long-lived animals share common changes while others appear to use opposing strategies. The datasets and analyses contained within and the user-friendly website developed will be of interest to researchers interested in complicated transcriptomic analyses and/or the biology of aging.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      This manuscript by Rudich ZD et al. systematically profiled the transcriptomic changes in nine long-lived C. elegans mutants and presented a careful and informative comparative analysis of these aging-related changes. In addition to these valuable datasets and bioinformatics analyses, the authors performed a large-scale RNAi screen to assess the role of the differentially expressed genes (DEGs) in these mutants and identify several potential targets to promote healthy aging. Moreover, the authors have provided a user-friendly website to examine genes of interest in those longevity mutants from their datasets.

      Strengths:

      Compared to previous transcriptomic analyses of these mutants in different reports, this study minimized the technical variations and benefitted from the advances in RNA-Seq technology and bioinformatics tools. Therefore, it should provide a more consistent and comprehensive view of the molecular mechanisms underlying the longevity of these mutants. The datasets in this manuscript are valuable to other researchers in the biology of aging.

      Weaknesses:

      Meanwhile, since these mutants have been extensively studied, the advance of this study in unknown ageing mechanisms remains limited.

      Comments on revised version.

      In the revised manuscript, the authors have addressed most of my concerns. In the text of this manuscript, the authors should still include more discussion on why osm-5 and daf-2 are categorized into two different groups.

    3. Reviewer #2 (Public review):

      Summary:

      In the manuscript titled "Multiple Molecular Pathways to Longevity: Opposing Gene Expression Programs Define Distinct Aging Strategies", the authors investigated diverse genetic pathways that contribute to lifespan extension in Caenorhabditis elegans and aimed to identify shared and distinct molecular mechanisms among various longevity mutants. Through comprehensive RNA sequencing of different longevity mutants representing seven distinct pathways, the authors showed that these mutants cluster into three primary groups based on their gene expression profiles. This transcriptomic analysis revealed that while some longevity genes are commonly regulated across multiple pathways, others exhibit opposing expression patterns, suggesting that distinct molecular strategies can lead to increased lifespan. Specifically, they identified a set of 196 genes that are consistently upregulated in most longevity mutants, many of which are involved in innate immunity and stress defense. By performing RNAi-based screening, the authors further validated the functional roles of several candidates, including C08F11.7, ugt-62, and K05C4.9, supporting their contributions to longevity and stress resistance. The authors conclude that longevity is mediated through multiple molecular pathways and provide a public online tool to study these complex transcriptomic landscapes.

      Significance:

      This study provides a systematic, side-by-side transcriptomic comparison of nine genetically distinct long-lived C. elegans mutants, revealing that lifespan extension arises from both shared and opposing gene expression programs. By identifying three distinct longevity groups and demonstrating that key pathways can be modulated in opposite directions to achieve long life, the work challenges the notion of a single universal transcriptional signature of aging. Importantly, functional validation shows that select commonly regulated genes can directly modulate lifespan and stress resistance, highlighting actionable molecular targets for promoting healthy aging.

      Comments on revised version:

      The authors addressed my concerns successfully.

    4. Author response:

      Reviewer #1:

      Major comments

      (1) Although I myself believe that the datasets in this study should be more consistent and comprehensive, the authors should perform a data mining analysis of previously reported transcriptomic changes of these mutants or similar mutants in the same longevity pathway and compare the reported changes with their findings to highlight the necessity and advances of this study.

      According to this suggestion, we have compared the differentially expressed genes identified in this study to previous gene expression studies involving these long-lived mutant strains. To our knowledge no previous studies have examined gene expression in sod-2 or ife-2 mutants, and at the time that we performed the RNA sequencing gene expression in osm-5 worms had not been examined (it took us a long time to complete this paper). We have included weighted Venn diagrams to illustrate the overlap and supplemental tables to list the overlapping di erentially expressed genes. For our current study, we felt it was important to compare RNA-seq data generated under exactly the same experimental and analysis paradigms in order to best compare across the nine long-lived mutants. These new analyses are included in Figures S19 – S25 and Table S2 . Please see lines 111-114, Figure S19-25, and Table S2.  

      (2) This manuscript does not perform any regulon or transcription factor (TF) analyses. TFs are the drivers of the transcriptomic changes and multiple conserved TFs (e.g., daf-16) have already been identified in these pathways. Therefore, it is necessary to examine and compare the regulons/TFs in these new datasets by bioinformatics. Such analyses can: a) provide more information of the driving force of these transcriptomic changes; b) show the role of these known longevity TFs; c) propose new TFs driving longevity; d) support the findings of 'longevity strategies' and 'longevity groups' from the perspective of TFs.

      According to this suggestion, we have now performed transcription factor analysis on the RNA-seq data to determine which transcription factors might be driving the longevity-associated transcriptional changes. To do this we used two complementary approaches: (1) transcription factor inference, which is based on the coordinated expression changes of known transcription factors; and (2) motif enrichment analysis, which is based on identifying transcription factor binding motifs in the promoters of di erentially expressed genes. After identifying which transcription factors were identified for each individual mutant, we then compared the identified transcription factors across all nine mutants. Interestingly, while 33 of the same transcription factors were implicated in group 1 and group 2 longevity mutants, 25 are modulated in different directions (activated in group 1, repressed in group 2 or vice versa) while only 5 are modulated in the same direction. This indicates that although group 1 and group 2 longevity mutants may modulate overlapping pathways to achieve long lifespan, in most cases these pathways are modulated in opposite directions. These new analyses are included in Figure S31 and Table S5. Please see lines 194-208, Figure S31, and Table S5.  

      (3) osm-5 and daf-2 are categorized into two different groups in this study. Since the longevity of cilia (-) mutants is through daf-16, the same master TF driving daf-2 longevity, please perform further analyses or discussion to clarify this issue.

      Loss of daf-16 is generally detrimental to lifespan. Disruption of daf-16 decreases the lifespan of all nine long-lived mutants that we examined (see supplemental table in our review paper PMID:37127095). However, loss of daf-16 also decreases wild-type lifespan. Thus, without further evidence it is hard to distinguish between the loss of daf-16 non-specifically decreasing lifespan verse activation of DAF-16 actually contributing to lifespan extension. In daf-2 mutants and the long-lived mitochondrial mutants there is increased nuclear localization of DAF-16 and upregulation of DAF-16 target genes. The differentially expressed genes in the long-lived mitochondrial mutants exhibit about a 50% overlap with the differentially expressed genes in daf-2 mutants (see Author response image 1). In contrast, osm-5 mutants show upregulation of some DAF-16 upregulated genes, no change in some DAF-16 upregulated genes and downregulation of other DAF-16 upregulated genes (see Author response image 1). Only about 10% of the differentially expressed genes in osm-5 mutants overlap with differentially expressed genes in daf-2 mutants. We believe that these results are consistent with loss of DAF-16 causing a general decrease in lifespan and not specifically contributing to osm-5 longevity. These comparisons will be included in a manuscript that we are currently preparing on osm-5 mutant longevity.

      Author response image 1.

      (4) This manuscript focused on genes whose RNAi suppressed the mutants longevity. Please also use bioinformatics to analyze the functions of those whose RNAi extends the mutants longevity, because these genes could tell the health price these mutants pay and help improve ageing interventions by reducing side effects.

      We perform enrichment analysis for both genes upregulated and downregulated in the long-lived mutant strains. The downregulated genes are involved in translation, ribosome biogenesis and gene expression. For the RNAi screen, we aimed to identify genes that are contributing to longevity and so we looked for a decrease in the lifespan of long-lived mutants when treated with RNAi. We did not screen for genes that extend the long-lived mutants longevity. While we did, nonetheless, identify multiple RNAi clones that increased either daf-2 or nuo-6 lifespan, there were not enough genes to identify any patterns of enrichment.

      (5) (OPTIONAL) I strongly suggest a comprehensive comparison of these transcriptomic changes in long-lived mutants with published age-related transcriptomic changes in wild type worms.

      According to this suggestion, we have now compared the differentially expressed genes that we identified in the nine long-lived mutants with genes that were found to be differentially expressed with aging. Interestingly, the group 2 long-lived mutants show a larger overlap for genes modulated in the opposite direction as aging (genes downregulated during aging are upregulated in eat-2 and osm-5 mutants). We have added this new analysis to our manuscript. Please see lines 210-223, Figure S32 and Table S6.

      Minor comments

      (1) Please further clarify the analysis of DEGs correlated with lifespan extension in Fig. 2 by a depiction. In Fig. 2C and D, please label data dots from different strains with different colors.

      According to this suggestion, each strain has been labelled a different colour.

      (2) In Fig. 3 and S20, please label the percentage of overlapping genes on top of each bars.

      We have now labelled the percentage of overlapping genes in Figure 3 and S20 (now S27).

      Reviewer #2:

      Major comments

      (1) While the authors identified a set of 196 upregulated genes, the rationale for narrowing these down to the three final candidates (C08F11.7, ugt-62, and K05C4.9) is not clearly described. The authors show that genetic inhibition of several genes, including DC2.5, C05B5.5, T07C4.5, and W03B1.7, decreases lifespan in both nuo-6 mutants and wild-type animals. However, the authors did not describe why these additional validated candidates, which also showed significant effects on longevity, were not pursued for further

      characterization. The authors should explicitly state the criteria used to prioritize these three genes over the other validated genes.

      Due to the costs and time involved in generating and characterizing new strains, we decided that we would select three strains to study further as a proof-of-principle. When deciding which genes to study further, we considered several approaches. In the end, we chose to use the strength/reproducibility of the increase in weighted mortality to identify genes with a clear, consistent impact. C05B5.5 and T07C4.5 were ruled out because they had an inconsistent impact on weighted mortality (Figure S28). W03B1.7 was ruled out because it did not have a strong enough e ect on weighted mortality (Figure S28). That narrowed it down to C08F11.7, ugt-62, DC2.5, and K05C4.9. Of those 4, C08F11.7, ugt-62, and K05C4.9 have the greatest consistent impact on weighted mortality (Figure S28) and so these genes were chosen. We have updated the manuscript to include this justification for focussing on C08F11.7, ugt-62, and K05C4.9. Please see lines 273-278.

      (2) The authors conclude that longevity can be mediated by multiple molecular pathways. However, it remains unclear whether these distinct strategies can operate simultaneously or are mutually exclusive. The authors need to test whether lifespan extension in a Group 1 mutant is further enhanced or suppressed by the knockdown of a key Group 2-specific genes. These experiments would help determine these pathways act additively, antagonistically, or as partially redundant survival programs.

      This is an excellent suggestion. While our data identify several genes that are regulated in opposite directions in group 1 and group 2 longevity mutants, we do not yet know the extent to which each of these genes contribute to the longevity of group 1 and group 2 mutants. The three genes that we focused on for further characterization (C08F11.7, ugt-62 and K05C4.9) are upregulated in group 1 longevity mutants but not group 2 mutants. Contrary to what might be expected, RNAi knockdown of these genes does not decrease the lifespan of the group 1 longevity mutant daf-2 but does decrease the lifespan of the group 2 longevity mutant eat-2. We recently reviewed the e ect of di erent resilience pathways on the lifespan of long-lived genetic mutants. Disruption of daf-16, sek-1, skn-1, hsf-1, ire-1 and trx-1 can decrease lifespan in both group 1 and group 2 longevity mutants, but also decreases lifespan in wild-type worms suggesting that at least in some mutants the e ect on longevity may be non-specific. Disruption of hif-1 does not a ect the longevity of group 2 mutants, but does a ect the lifespan of some group 1 mutants (clk-1, isp-1, nuo-6) but not others (daf-2, glp-1). To more definitively answer the question, it would be interesting to cross different combinations of group 1 and group 2 longevity mutants to see the extent to which different longevity groups synergize. This is something we are currently working on for a separate manuscript. We have added these points to the revised manuscript. Please see lines 363-381.

      (3) The authors provide interesting data on overexpression of the three candidate genes. However, whereas C08F11.7 clearly demonstrates both necessity and sufficiency for lifespan extension, overexpression of ugt-62 and K05C4.9 does not independently extend lifespan. To strengthen the manuscript, the authors should expand the discussion of these divergent results and clarify possible explanations.

      According to this suggestion, we have expanded our discussion to discuss possibilities of why these genes might be having different effects on lifespan. Please see lines 411-423.

      (4) Key citations are missing and the authors should add multiple citations including the following ones. Please cite the following paper and discuss the authors' finding with respect to the related work (Lee et al PMID: 40814218). Add citations in the sentence describing changes in the transcriptome of C. elegans associated with age (Lee et al., PMID: 38508494). Furthermore, please cite papers describing the overviews of survival assay using C. elegans (Kwon et al., PMID: 40436148, Hwang et al., PMID: 40436147).

      We have added the suggested citations to the revised manuscript. Please see lines 211 (Ref #39), 307 (Ref #41), 423 (Ref #54) and 436 (Ref #55).

      Minor comments

      (1) To improve readability, please provide the full names for all abbreviations at their first appearance in the manuscript.

      We have added the full names for each abbreviation on first appearance.

      (2) Please ensure that the labels in the figures match the text exactly. For instance, if different promoters are used for generating overexpression animals, it may be helpful to indicate the specific promoter in the figure panel or legend for clarity.

      We have ensured that the nomenclature in the text and the figures is the same. We have noted the promoter used for the overexpression strains in the figure legend.

      (3) For all lifespan and stress resistance assays, please include the total number of animals (n) and the number of independent biological replicates (N) in the figure legends to confirm statistical reliability.

      We have added the number of animals and independent biological replicates to the figures and figure legends.

      (4) Please clearly specify the exact developmental stage of the animals used for the survival assays in the Materials and Methods section.

      We have updated the methods to describe the developmental stages used for the survival assays.

    1. eLife Assessment

      This valuable study provides insights into the role of MATR3 in oocyte maturation and folliculogenesis, using conditional knockout mice and in vitro follicle culture systems to show that MATR3 is required for oocyte growth and gene transcription, with downstream effects on follicle development. The evidence is solid, but some minor inadequacies in replication of key methods and independent validation reduce confidence in the conclusions. The work will be of interest to researchers in reproductive biology and fertility.

    2. Reviewer #1 (Public review):

      Summary:

      This study aims to clarify MATR3's function and molecular mechanism in oocyte growth and maturation, explore its association with OMA and its potential as a diagnostic and therapeutic target using specific knockout mouse models, human OMA samples and multi-omics technologies. And it has fully achieved preset objectives with results strongly supporting conclusions. Specifically, it addresses the gap in the synergistic mechanism of epigenetic and secretory signals regulated by RNA-binding proteins (RBPs) in oocyte growth and enriches the molecular etiological spectrum of oocyte maturation disorders. It is the first time to reveal the conservative function of MATR3 in multiple species, providing a paradigm for cross-species research on RBPs in the field of reproductive biology. And it provides a new candidate target for OMA, a clinically refractory infertility disease, and is expected to promote the optimization of assisted reproductive technology and the development of precision medicine.

      Strengths:

      The strengths of this study are significant and prominent. First, the research system is comprehensive, integrating knockout mouse models, in vitro knockdown models, multi-species (mouse, porcine and human) verification, combined with scRNA-seq, LACE-seq, CO-IP and other multi-omics and molecular biology technologies, forming a complete and progressive evidence chain. Second, the mechanism analysis is in-depth, clarifying the dual molecular mechanisms of MATR3 regulating the transcriptional synthesis and secretion of GDF9 through "recruiting KDM3B to regulate H3K9me2 demethylation" and "directly binding to Rdx mRNA", with a clear logical closed loop. Third, the clinical correlation is close. It is the first time to find abnormal nuclear localization of MATR3 in oocytes of OMA patients, providing new clues for clinical disease mechanism research, and verifying the downstream function of GDF9 through rescue experiments, effectively enhancing the translational value of the results.

      Weaknesses:

      This study included only one OMA patient's oocyte sample. Without clinical screening for MATR3 mutations or abnormal expression, establishing a causal relationship between MATR3 and OMA remains difficult.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout mouse models together with in vitro follicle culture and molecular analyses. The authors aim to determine whether MATR3 regulates oocyte maturation and follicle development and to explore potential mechanisms linking MATR3 function to transcriptional and epigenetic regulation in growing oocytes.

      Strengths:

      A major strength of the work is the use of a conditional knockout mouse model combined with complementary in vitro follicle culture approaches, which together provide a useful framework for examining gene function during oocyte development. The study also attempts to integrate cellular phenotypes with molecular analyses of transcriptional activity and epigenetic markers.

      Weaknesses:

      Several weaknesses limit the strength of the conclusions. These include insufficient validation of key experimental manipulations (such as the efficiency of MATR3 knockdown in siRNA experiments), limited quantification or statistical analysis for some datasets, inconsistencies between the text and presented data in certain figures, and incomplete methodological descriptions that make it difficult to fully evaluate reproducibility.

      Comments on revised version.

      Thank you for submitting the revised manuscript. I believe the revisions have substantially improved the quality and clarity of the study, and the authors have addressed the major concerns raised during the initial review.

    4. Reviewer #3 (Public review):

      Summary:

      The study aims to elucidate the dual molecular mechanisms of the RNA-binding protein MATR3 in oocyte growth and maturation. The authors propose that MATR3, highly expressed in growing oocytes (GOs), regulates oocyte quality through two pathways: epigenetically, by recruiting KDM3B to remove the repressive H3K9me2 mark at the Gdf9 locus to activate transcription; and post-transcriptionally, by binding Rdx mRNA to maintain microvillus structure for GDF9 secretion. This mechanism ensures oocyte-granulosa cell communication and female fertility. The study also explores the link between MATR3 and human oocyte maturation arrest (OMA).

      Strengths:

      The study proposes an innovative dual-mechanism model encompassing "epigenetic transcriptional activation and cytoskeletal regulation," which not only expands the functional understanding of RNA-binding proteins in chromatin regulation but also reveals the coordination between nuclear transcription and organelle structure. By integrating scRNA-seq and LACE-seq, the authors constructed a comprehensive regulatory network for MATR3, identifying both key targets and numerous potential molecules, thereby providing rich resources for future mechanistic studies. Furthermore, the inclusion of oocyte samples from human OMA patients directly links the basic findings to clinical reproductive disorders. Despite the limited sample size, this approach demonstrates strong translational potential.

      Weaknesses:

      The partial phenotypic improvement achieved by exogenous GDF9 supplementation suggests that the downstream effector pathways may involve a more complex network regulation, implying that the current interpretation of GDF9 central role could be further explored. Regarding the developmental abnormalities of granulosa cells in the conditional knockout model, their pathological origins require in-depth analysis to determine whether they represent primary alterations or secondary adaptive responses resulting from the loss of oocyte signaling.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study aims to clarify MATR3's function and molecular mechanism in oocyte growth and maturation, explore its association with OMA, and its potential as a diagnostic and therapeutic target using specific knockout mouse models, human OMA samples, and multi-omics technologies. And it has fully achieved preset objectives with results strongly supporting conclusions. Specifically, it addresses the gap in the synergistic mechanism of epigenetic and secretory signals regulated by RNA-binding proteins (RBPs) in oocyte growth and enriches the molecular etiological spectrum of oocyte maturation disorders. It is the first time the conservative function of MATR3 has been revealed in multiple species, providing a paradigm for cross-species research on RBPs in the field of reproductive biology. It also provides a new candidate target for OMA, a clinically refractory infertility disease, and is expected to promote the optimization of assisted reproductive technology and the development of precision medicine.

      Strengths:

      The strengths of this study are significant and prominent. First, the research system is comprehensive, integrating knockout mouse models, in vitro knockdown models, multi-species (mouse, porcine, and human) verification, combined with scRNA-seq, LACE-seq, CO-IP, and other multi-omics and molecular biology technologies, forming a complete and progressive evidence chain. Second, the mechanism analysis is in-depth, clarifying the dual molecular mechanisms of MATR3 regulating the transcriptional synthesis and secretion of GDF9 through "recruiting KDM3B to regulate H3K9me2 demethylation" and "directly binding to Rdx mRNA", with a clear logical closed loop. Third, the clinical correlation is close. It is the first time to find abnormal nuclear localization of MATR3 in oocytes of OMA patients, providing new clues for clinical disease mechanism research, and verifying the downstream function of GDF9 through rescue experiments, effectively enhancing the translational value of the results.

      Weaknesses:

      This study included only one OMA patient's oocyte sample. Without clinical screening for MATR3 mutations or abnormal expression, establishing a causal relationship between MATR3 and OMA remains difficult.

      We greatly appreciate positive comments and constructive feedback on our manuscript.

      We are encouraged that you recognize the novelty, rigour, and clinical relevance of our study on MATR3 in oocyte development and OMA. We have carefully considered your comments and revised the manuscript accordingly. We will further expand the OMA patient cohort in future studies to verify the causal relationship between MATR3 and OMA.

      Reviewer #2 (Public review):

      Summary:

      This study investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout mouse models together with in vitro follicle culture and molecular analyses. The authors aim to determine whether MATR3 regulates oocyte maturation and follicle development and to explore potential mechanisms linking MATR3 function to transcriptional and epigenetic regulation in growing oocytes.

      Strengths:

      A major strength of the work is the use of a conditional knockout mouse model combined with complementary in vitro follicle culture approaches, which together provide a useful framework for examining gene function during oocyte development. The study also attempts to integrate cellular phenotypes with molecular analyses of transcriptional activity and epigenetic markers.

      Weaknesses:

      Several weaknesses limit the strength of the conclusions. These include insufficient validation of key experimental manipulations (such as the efficiency of MATR3 knockdown in siRNA experiments), limited quantification or statistical analysis for some datasets, inconsistencies between the text and presented data in certain figures, and incomplete methodological descriptions that make it difficult to fully evaluate reproducibility.

      We greatly appreciate your constructive comments and suggestions. We are grateful for the recognition of our conditional knockout mouse model and experimental design. We have carefully addressed all the weaknesses mentioned by the reviewer, including the validation of key experiments, quantitative and statistical analysis, consistency between text and figures, and detailed methodological descriptions. Details are described point-by-point below.

      Reviewer #3 (Public review):

      Summary:

      The study aims to elucidate the dual molecular mechanisms of the RNA-binding protein MATR3 in oocyte growth and maturation. The authors propose that MATR3, highly expressed in growing oocytes (GOs), regulates oocyte quality through two pathways: epigenetically, by recruiting KDM3B to remove the repressive H3K9me2 mark at the Gdf9 locus to activate transcription; and post-transcriptionally, by binding Rdx mRNA to maintain microvillus structure for GDF9 secretion. This mechanism ensures oocyte-granulosa cell communication and female fertility. The study also explores the link between MATR3 and human oocyte maturation arrest (OMA).

      Strengths:

      The study proposes an innovative dual-mechanism model encompassing "epigenetic transcriptional activation and cytoskeletal regulation," which not only expands the functional understanding of RNA-binding proteins in chromatin regulation but also reveals the coordination between nuclear transcription and organelle structure. By integrating scRNA-seq and LACE-seq, the authors constructed a comprehensive regulatory network for MATR3, identifying both key targets and numerous potential molecules, thereby providing rich resources for future mechanistic studies. Furthermore, the inclusion of oocyte samples from human OMA patients directly links the basic findings to clinical reproductive disorders. Despite the limited sample size, this approach demonstrates strong translational potential.

      Weaknesses:

      The partial phenotypic improvement achieved by exogenous GDF9 supplementation suggests that the downstream effector pathways may involve a more complex network regulation, implying that the current interpretation of GDF9's central role could be further explored. Regarding the developmental abnormalities of granulosa cells in the conditional knockout model, their pathological origins require in-depth analysis to determine whether they represent primary alterations or secondary adaptive responses resulting from the loss of oocyte signaling.

      We greatly appreciate your positive and insightful comments on our study. We are grateful for the recognition of our novel dual-mechanism model, comprehensive multi-omics analysis, and translational potential from basic research to clinical OMA. We have carefully addressed the weaknesses raised by the reviewer, including in-depth discussion of the GDF9-centered regulatory network and clarification of the origin of granulosa cell abnormalities. More details point-by-point responses are provided below.

      Recommendations for the authors:

      Point-by-point responses to reviewers’ comments

      We thank the reviewer very much for his/her reviewing of our work, and we appreciate the constructive comments and suggestions that have helped us to prepare an improved revision. Based on the comments of the reviewer, we have carefully revised the manuscript by performing some new experiments.

      Reviewer #1 (Recommendations for the authors):

      (1) Did most of the follicles cultured in vitro reach the antral follicle stage after 6 days?

      We greatly appreciate your insightful question. We statistically analyzed the survival rate and antral follicle ratio of in vitro-cultured follicles after 6 days of culture. Due to differences in culture systems and protocols, the follicle survival rate in our study (57.43 ± 3.11%) was different from that reported in previous literature (92 ± 10%). However, the proportion of antral follicles among surviving follicles was highly consistent between our results and published data (83 ± 13% vs 80.87 ± 3.27%) (Cortvrindt and Smitz 2002).

      Author response image 1.

      Ratio and survival rate of antral follicles after 6 days of culture. n = 3. Data are represented as mean ± SD.

      (2) In Figure 2F, at which stage did MATR3 begin to affect oocyte diameter?

      Thank you for your careful observation. Our morphological analysis of oocytes collected from PD14 and PD23 mice showed no significant difference in oocyte diameter between the cKO and Ctrl groups at the GO stage (Fig. S3D, E). However, oocytes in the cKO group became significantly smaller than those in the Ctrl group once they reached the FGO stage (Fig. 2E, F). Taken together, these results indicate that the growth defect caused by MATR3 deletion begins to manifest during the transition from the GO to FGO stage, with significant reduction in oocyte diameter clearly observed at the FGO stage as shown in Figure 2F.

      (3) What was the developmental potential of oocytes in Matr3-knockout mice?

      Thank you for this important question. Compared with the Ctrl group, oocytes derived from cKO mice showed a drastically reduced fertilization rate (91.55 ± 1.96% vs 10.55 ± 4.78%) and almost completely failed to develop to the blastocyst stage (76.62 ± 7.56% vs 3.67 ± 3.38%). These results clearly demonstrate that maternal deletion of Matr3 severely compromises the developmental potential of mouse oocytes, including fertilization capacity and subsequent early embryonic development.

      Author response image 2.

      Results of in vitro fertilization of oocytes. 2-cell: 2 days after fertilization; blastocyst: 4 days after fertilization. Data are represented as mean ± SD. ***P < 0.001.

      (4) The legend labels in the figures should not be bold.

      Thank you for your valuable suggestion. We have revised all the figures accordingly.

      Reviewer #2 (Recommendations for the authors):

      This manuscript investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout (cKO) mouse models combined with in vitro follicle culture approaches. The topic is relevant to the field of reproductive biology and provides potentially important insights into the molecular mechanisms regulating oocyte maturation and follicle development.

      While the study presents interesting observations and utilizes both in vivo and in vitro experimental systems, several issues need to be addressed before the manuscript can meet the expected standards. These include concerns related to data interpretation, validation of experimental approaches, completeness of methodological descriptions, and clarity in data presentation. In addition, the manuscript requires substantial language editing to improve clarity and readability.

      The comments below outline major issues that should be addressed to strengthen the manuscript, as well as specific minor points regarding presentation and clarity.

      (1) The manuscript requires substantial revision to improve the written language and grammar. Numerous sentences are unclear or awkwardly phrased, which makes interpretation of the results difficult in several sections. The authors are strongly encouraged to have the manuscript professionally edited or thoroughly revised for language and clarity before making a resubmission.

      We sincerely appreciate the careful and constructive comments on the language quality and clarity of the manuscript. We fully agree that the written language, grammar, and sentence structure need substantial improvement to ensure the results are presented clearly and accurately.

      To address these concerns thoroughly, we have carefully revised the entire manuscript, including correcting grammatical errors, refining awkward phrasing, and restructuring unclear sentences to enhance readability and logical flow. In addition, we have sought professional language editing support to further polish the English expression and ensure the manuscript meets the linguistic standards of the journal.

      All revisions related to language and clarity have been completed, and we believe the revised version is significantly improved in terms of readability and precision.

      (2) Interpretation of oocyte maturation results (Line 140; Figure 2E, H). The manuscript states: "During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig. 2E, H)." However, Figure 2H appears to show that a small proportion of knockout oocytes do reach the MII stage. Therefore, the description in the text seems inconsistent with the data presented. The authors should clarify the exact maturation rates in both groups, revise the text to accurately reflect the data, and provide statistical analysis to support the stated conclusions.

      Thank you for your valuable comment. A small proportion of knockout oocytes from PD23 cKO mice can indeed develop to the MII stage. We have revised the corresponding description and supplemented the statistical analysis of maturation rates to support our conclusion.

      Line 140: “During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig.2E, H).” have been replaced by “During in vitro maturation, the proportion of oocytes from PD23 cKO mice developing to metaphase II stage was significantly reduced (Fig.2E, H, 54.9±2.08% vs 9.57±1.11%).”

      (3) Human oocyte sample size: In Figure 1D, it is unclear how many human oocytes were analyzed. It is important to specify the sample size (n) for all experiments. The authors should clearly indicate the number of oocytes analyzed in this experiment. Provide this information either in the figure legend or in the main text.

      Thank you for this important comment. We agree that the altered subcellular localization of MATR3 in human OMA oocytes is of great physiological significance for understanding the functional role of MATR3 during oocyte development.

      Unfortunately, during a 3‑month period of sample collection, we examined MATR3 localization in immature oocytes that failed to reach the MII stage, obtained from 11 women undergoing IVF treatment. Among these samples, only one donor’s oocytes exhibited the NSN chromatin configuration. Excitingly, these NSN‑stage oocytes from this donor clearly showed the loss of MATR3 nuclear localization, which strongly supports the critical role of MATR3 during oocyte growth and maturation. We have now clearly stated the sample size (n = 11) in the figure legend and main text as suggested. In future studies, we will continue to collect more human oocyte samples to further validate these observations with an expanded sample size.

      (4) Figure annotation issue: The figure legend for Figure 1F refers to an arrow, but no arrow is visible in the figure panel. Please correct this inconsistency by either adding the appropriate arrow to the figure or revising the legend accordingly.

      Thank you for pointing out this error. We have revised the figure legend for Figure 1F accordingly to correct this inconsistency.

      (5) Description of follicle analysis (Line 146): The sentence: "This was reinforced by the data of available follicles within the follicles of mice on PD35 (Fig. 2I, J)." is incorrect or poorly phrased. It should likely read: "...available follicles within the ovaries of mice at PD35...".

      We really appreciate your constructive suggestion on the phrasing. We have revised this sentence in the revised manuscript accordingly.

      (6) Quantification of proliferating cells: Figure 2K shows Ki-positive cells, but quantitative analysis is not provided. The authors should quantify the number or proportion of Ki-positive cells in both control and cKO groups and include statistical analysis to support any claims regarding differences in proliferation.

      Thank you for your valuable suggestion. We have quantified the number of Ki‑67‑positive granulosa cells in both control and cKO groups and performed the corresponding statistical analysis.

      The quantitative results have been added to Fig. S3G, and the relevant description has been supplemented in the main text at Line 150 to support our conclusion regarding cell proliferation differences.

      Line 150: “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K) in cKO mice were lower than those found in the Ctrl.” have been replaced by “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K, Fig. S3G) in cKO mice were lower than those found in the Ctrl (73.55±13.29% vs 24.65±7.80%).”

      (7) Validation of findings in the in vivo cKO model (Figure 3): The development of an in vitro follicle culture system is an interesting and valuable component of the study. However, several key analyses performed in vitro (e.g., transcription assays and analysis of epigenetic markers) should ideally also be validated in oocytes derived from the in vivo cKO model.

      Thanks for the valuable concern. We fully agree with you that the in vivo cKO model should be used to validate several key analyses performed in vitro. We collected growing oocytes from Ctrl and cKO mice and conducted transcription assays as well as analysis of epigenetic markers. The results showed that Matr3 knockout significantly downregulated transcriptional activity in GO and increased H3K9me2 levels (Author response image 3), which is consistent with our in vitro findings (Fig 3D E I J). These in vivo results confirm that MATR3 plays a critical role in regulating GO transcriptional activity and H3K9me2 levels.

      Author response image 3.

      Matr3 knockout results in the reduction of transcriptional activity. A EU staining (green) in GO collected from Ctrl and cKO. n = 15. B Quantification of the mean fluorescence intensity of EU in oocytes. C H3K9me2 staining (red) in GO collected from Ctrl and cKO. n = 15. D Quantification of the mean fluorescence intensity of H3K9me2 in oocytes. Scale bar: 20 μm. Data are represented as mean ± S.D. ***P < 0.001.

      (8) To strengthen the conclusions, the authors should consider repeating key experiments using oocytes directly isolated from the cKO mice. This would help confirm that the observed effects are not artifacts of the in vitro culture system.

      Thank you for this valuable and constructive suggestion. We fully agree that the conditional knockout mouse model is essential for verifying the physiological significance of MATR3 in vivo.

      To address this point, we have validated multiple key in vitro findings using oocytes directly isolated from cKO mice. For instances, the changes in oocyte transcriptional activity (EU staining) (in Comments 7), H3K9me2 levels (in Comments 7), GDF9 levels (Fig 4.B C E F), and OO-Mvi (Fig 6.A B, Author response image 4) all showed consistent trends with our in vitro knockdown results. In addition, the complete infertility phenotype of cKO female mice further demonstrates that MATR3 is indispensable for oocyte growth and meiotic maturation. We have also provided supplemental data from GDF9 rescue experiments and sequencing analysis performed in the mouse model.

      Author response image 4.

      Matr3 knockdown impairs the structural integrity of oocyte OO-MVi. A p-ERM staining (green) showing the OO-Mvi in oocyte from NC and si-Matr3. B Quantification of the number of Oo-Mvi vesicles (n = 6). Scale bar: 20 μm. Data are represented as mean ± SD. ***P < 0.001.

      In conclusion, the core conclusions of this study are supported by the mutual validation of key experimental results from MATR3-specific knockdown in vitro and Matr3 conditional knockout mouse models in vivo.

      (9) (1) Validation of MATR3 knockdown: The in vitro MATR3 knockdown experiment presented in Figure 4G raises an important concern: it is unclear whether Matr3 knockdown was effectively achieved in the oocytes analyzed. The authors should provide direct evidence of knockdown efficiency, for example, immunostaining for MATR3 protein on the oocytes. Without such validation, it is difficult to interpret the functional outcomes observed.

      As requested, we have provided direct evidences of the knockdown efficiency via immunostaining, which is now presented in Fig.S2D.

      In this experiment, oocytes from early growing follicles (approximately 150 μm in diameter) were microinjected with Matr3 siRNA. Following 5 days of continuous in vitro culture, oocytes from both the NC and si-Matr3 groups were isolated and subjected to immunofluorescence staining to assess protein levels. As shown in the figure, the oocytes at this stage exhibited the characteristic non-surrounded nucleolus (NSN) chromatin configuration. We observed robust MATR3 protein expression within the nucleus of NC oocytes, whereas the MATR3 protein levels were markedly reduced in the si-Matr3 group. These results confirm the successful construction of the Matr3 knockdown model in early growing follicle oocytes.

      (9) (2) Furthermore, it would be more convincing if the authors could perform the Gdf9 supplementation experiments using follicles isolated from the cKO mice, rather than relying solely on siRNA knockdown in vitro. Such experiments would provide clearer and more physiologically relevant evidence. If these experiments were attempted but did not produce similar results, this should be discussed.

      Thank you for your valuable and insightful suggestion. We fully agree that performing GDF9 supplementation experiments using follicles isolated from cKO mice would provide more direct and physiologically relevant evidence to strengthen our conclusions.

      Unfortunately, when we attempted to conduct GDF9 rescue experiments on follicles from cKO mice, neither the control nor cKO follicles were able to develop to the antral follicle stage (n=3). We speculate that this was caused by insufficient bioactivity of the veterinary-grade FSH been used, as compared to the imported FSH been provided by NHPP. Unfortunately, this particular FSH product has been discontinued. We are currently actively seeking and attempting to purchase new, qualified FSH reagents to repeat these experiments and further validate our findings in future work.

      (10) Figure citation order: Figures are not cited sequentially in the text. For example, Figure 6J is described first (line 250), followed by Figure 6A. Figures should be discussed in logical order, typically starting from panel A. Please revise the text to ensure that figure panels are introduced sequentially.

      Thank you for this careful and important comment.

      We have carefully revised the citation order of all figure panels in the main text, especially for Figure 6, to ensure they are introduced sequentially from panel A to the last panel in logical and numerical order, rather than being cited out of sequence.

      The corresponding adjustments have been made in the revised version of the manuscript.

      (11) Figure 6J interpretation: The purpose of the images shown in this figure is unclear. The authors should provide higher magnification images to clearly visualize the Oo-Mvi structures and include quantification of the observed phenotype to support the interpretation. It is important because the main findings of the paper heavily rely on these results.

      Thank you for this valuable and constructive suggestion. We fully agree that higher‑magnification images and quantitative analysis are essential to clearly demonstrate the Oo‑Mvi structures and reliably support our conclusions, especially given the importance of these results to the main findings of this study.

      Accordingly, we have replaced the original panels in Figure 6A with higher‑magnification images to better visualize Oo‑Mvi structures. In addition, we have supplemented the corresponding quantitative analysis of the observed phenotype to strengthen the interpretation of this figure (Fig 6B). All revisions have been incorporated into the revised manuscript.

      (12) Incomplete Materials and Methods section: The Materials and Methods section lacks important experimental details required for reproducibility. Specifically, the Matr3 flox mouse model. Either provide the appropriate reference describing the Matr3 floxed mice or include details on how the floxed allele was generated.

      Thank you for your valuable and careful comment. We fully agree that detailed experimental information in the Materials and Methods section is crucial for ensuring the reproducibility of the study. And we apologize for the omission of key details regarding the Matr3 flox mouse model.

      In response to your suggestion, we have thoroughly supplemented the relevant experimental details in the Materials and Methods section of the revised manuscript, including the specific construction strategy of the Matr3 floxed allele. These detailed descriptions will enable other researchers to reproduce our mouse model and verify the experimental results.

      All supplementary information has been integrated into the revised manuscript to meet the requirements of experimental reproducibility. We greatly appreciate your guidance in helping us improve the completeness and rigour of our study.

      (13) Follicle isolation: The manuscript does not describe how growing follicles were isolated. Please specify whether follicles were isolated using enzymatic digestion or mechanical dissection and provide sufficient methodological detail so that other researchers can reproduce the experiments.

      Thank you for this valuable comment. We agree that adding this information is essential for ensuring the reproducibility of our experiments. We have supplemented the corresponding description in the Materials and Methods section. Briefly, growing follicles were isolated by mechanical dissection using insulin syringes under a stereomicroscope, without any enzymatic digestion.

      Reviewer #3 (Recommendations for the authors):

      (1) Since KDM3B and MATR3 interact in cell lines, does this relationship affect the functional localization of KDM3B within oocytes? Specifically, does the localization of KDM3B change in cKO mice (e.g., nuclear export or aggregation)?

      Thank you for this insightful and constructive question. To address whether the interaction between KDM3B and MATR3 influences the functional localization of KDM3B in oocytes, we performed immunofluorescence staining to examine the subcellular distribution of KDM3B in cKO oocytes. Our results demonstrated that the nuclear localization of KDM3B remained unaltered; no obvious nuclear export or abnormal aggregation was observed in MATR3-deficient oocytes (Author response image 5).

      Based on these observations combined with our other experimental data, we propose that MATR3 regulates oocyte transcriptional activity through its physical interaction with KDM3B, rather than by controlling the nuclear targeting of KDM3B. Notably, despite unchanged nuclear localization of KDM3B in MATR3 cKO oocytes, we detected significantly elevated global levels of H3K9me2 and markedly reduced transcriptional activity (in Comments 7). These findings indicate that KDM3B loses its physiological function of demethylating H3K9me2 and promoting transcription in the absence of MATR3.

      In line with this mechanism, previous results showed that KDM3B knockout in female mice leads to follicle arrest at the secondary follicle stage and consequent infertility (Liu et al. 2015). Collectively, we conclude that in MATR3 cKO oocytes, although KDM3B is properly localized in the nucleus, it fails to execute its H3K9me2 demethylase activity, thereby impairing normal transcriptional regulation during oocyte development.

      Author response image 5.

      Matr3 knockout has no effect on KDM3B localization. KDM3B staining (red) in GO collected from Ctrl and cKO. n   = 50. Scale bar: 40 μm.

      (2) As the GDF9 rescue experiment only partially restores the phenotype, it is suggested to select 2-3 novel targets with high binding intensity and significant expression changes from the LACE-seq data (e.g., Igf2bp2 or Ccnb1 mentioned in the text) to further illustrate that MATR3 regulates a network.

      Thanks for this meaningful and constructive suggestion.

      We have supplemented the relevant data in Fig. S9 and further elaborated on these findings in the Discussion section in our revised manuscript, following your advice. Briefly, we collected growing oocytes from Ctrl and cKO mice and performed RT‑qPCR analysis to verify the expression of Igf2bp2 and Ccnb1 - two representative novel targets with strong binding intensity and significant expression changes identified from our LACE‑seq bioinformatics analysis. The results showed that both genes were significantly downregulated in cKO oocytes compared with controls, supporting the notion that MATR3 regulates a functional RNA network during oocyte development.

      (3) Are the granulosa cell defects primary or secondary? It is recommended to collect ovaries from earlier-stage cKO mice (e.g., PD7 or PD10) to examine the levels of FOXL2 and PCNA in granulosa cells.

      Thank you for your valuable comment. To clarify whether the granulosa cell defects are primary or secondary, we further investigated the temporal effect of MATR3 deficiency on granulosa cells by collecting ovaries from PD7, which is a critical period for primordial follicle activation and early follicular development.

      To evaluate the status of granulosa cells, we performed immunofluorescence staining on PD7 ovarian sections using FOXL2 (a specific marker for granulosa cells) to quantify the number of granulosa cells, and Ki67 (a proliferation-related marker) to assess the proliferative capacity of granulosa cells. The results, as shown in Author response image 6, demonstrated that there were no significant differences in either the number of granulosa cells or their proliferation levels in primary follicles between cKO and Ctrl.

      These findings are consistent with the data in our supplementary Fig S3F, where we observed no significant differences in the number of primordial follicles and growing follicles at PD7 between the two groups. Collectively, these results indicate that the activation of primordial follicles is not affected by MATR3 deficiency in oocytes, and the impairment of granulosa cells caused by oocyte-specific Matr3 knockout occurs at the secondary follicle stage rather than the early stages of follicular development. Therefore, we conclude that the granulosa cell defects in cKO mice are secondary to the oocyte dysfunction induced by MATR3 deficiency.

      Author response image 6.

      Loss of MATR3 in oocytes does not affect the number and proliferation of granulosa cells in primordial follicles. A Immunohistochemistry results showing granulosa cells in PD7 ovaries from Ctrl and cKO. B Quantification of granulosa cell number in the largest cross-section of primary follicles. C Quantification of the proliferation rate of granulosa cells in the largest cross-section of primary follicles. n = 15. Data are represented as mean ± SD. n.s., not significant.

      References:

      (1) Cortvrindt RG, Smitz JE. 2002. Follicle culture in reproductive toxicology: a tool for in-vitro testing of ovarian function? Human reproduction update 8: 243-254.

      (2) Liu Z, Chen X, Zhou S, Liao L, Jiang R, Xu J. 2015. The histone H3K9 demethylase Kdm3b is required for somatic growth and female reproductive function. International journal of biological sciences 11: 494-507.

    1. eLife Assessment

      This study reports important and invaluable findings that advance understanding of how attention is distributed between what we look at directly and what lies outside the center of gaze during active visual search. The evidence supporting the main claims is solid, with a large and rich dataset spanning multiple brain areas, although some aspects of the interpretation would benefit from additional controls and clearer separation of attention from eye-movement planning. The work will be of particular interest to researchers studying attention, visual perception, and eye movements behavior.

    2. Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye movement mediated search neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. Detailed association of simultaneously obtained eye movement sequences and neural parameters are well done. These are valuable data which will contribute to our understanding of attentional modulation in visual search.

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from key mid-tier (V4) and higher order (IT, PFC) areas. They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, marked by a high degree of feature and categorical specificity. That is, while attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This provides valuable data for the concept of a foveal-peripheral spatiotemporal attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors and looks towards and away from the target) and statistical rigor make these findings compelling. There will likely be additional future impacts of this study. For example, the eye movement patterns collected in this study may also provide a valuable dataset for future study of understanding search strategies. Goal-directed vs non-goal-directed task comparisons could be designed to test possible circuit models. Although much remains unknown regarding how and where frontal and temporal signals are integrated during active search, these data contribute important guideposts for future models of active visual search.

    3. Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Fig. 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provide a dataset that is well suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Fig. 2). As a result, the reported attentional modulation coincides with preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19) therefore likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Fig. S3) partially mitigate this concern by demonstrating that feature-based modulation persists through saccade execution.

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Fig. 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      [Editors' note: the authors have provided responses to each of these points.]

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to differentiate between foveal and peripheral attentional mechanisms in visual and frontal brain regions in monkeys engaged in a free-gaze visual search task.

      Strengths:

      The manuscript is clearly written, the question is important, and the behavioral task is interesting.

      Weaknesses:

      I have two major concerns.

      (1) The authors interpret divergence in neural responses to target vs nontarget as attention. But it is not. The subject has to attend to both target and nontarget stimuli to determine the stimulus category and thereby decide on the next action. Thus, divergence between target and nontarget responses could reflect categorical discrimination, but I am not sure this can be interpreted as attentional modulation. While it may be tempting to suggest that finding a stimulus of a specific category is "feature attention", analogous to, e.g., attending to the red stimulus, I don't believe this is correct. For the former, the animals have to attend to a stimulus, and examine the stimulus to determine the stimulus category, unlike a simpler discrimination, which may pop out. Given this, I am unconvinced that the interpretations in this manuscript are valid.

      We thank the reviewer for raising this concern. Selective attention is a process of focusing on goal-relevant stimuli (targets) while ignoring irrelevant distractions. Importantly, attentional selection is not limited to simple visual features (e.g., color, shape, or motion); it can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [1, 2], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [3, 4]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [5-8], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [7].

      Similarly, in our study, monkeys were trained to search for images that matched the category of the cue. The neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors. We also included only neural responses occurring prior to fixations associated with target selection, that is, before the monkeys made a behavioral choice, thereby controlling for potential contributions of target detection or decision-related signals to the observed effects.

      We have clarified and addressed this point in the Discussion as follows:

      “Feature-based attention to simple visual features such as color, shape, or motion has been extensively studied [1, 3, 5, 7-9, 11, 12, 64]. Attention can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [65, 66], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [6, 67]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [68-71], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [70]. In this study, the neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors.”

      (2) Regarding the RF classification of foveal and peripheral RFs for IT and PFC, prior work suggests that neurons in IT cortex (especially AIT) and PFC have RFs that largely include the foveal visual field. So, it would be important to include figures that show the RFs of neurons classified as foveal versus peripheral for all three areas.

      We thank the reviewer for raising this important point. We agree with the reviewer that neurons in IT cortex and PFC often have RFs that include the foveal visual field. We did record foveal units with both focal and broad foveal RFs; however, in our analysis we only included neurons with focal foveal RFs to exclude the influence of peripheral stimuli. We defined focal foveal-RF units as those that responded solely to the cue in the foveal region and not to items in the search array presented at least 5° away from the central fixation point, ensuring that their RFs did not extend to these peripheral locations. The items were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their RFs during fixations. By definition, their RFs were restricted to the central point. This is further supported by Fig. S1A-H, which shows no responses to items in the search array at peripheral locations. We have made modifications in the Results and Methods as follows:

      “Notably, the items in the search array were presented at least 5° from the central fixation point and were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their foveal RFs during fixations.”

      And:

      “In this study, our focus was on units with focal foveal RFs and units with localized peripheral RFs. All further analyses were conducted on these units.”

      We modified Fig. 1 to illustrate the RFs of neurons classified as peripheral, which were also characterized in our previous study using the same dataset [9]. The peripheral population exhibits no responses to the central cue (Fig. S1I–T).

      Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at the center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature-selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with a covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye-movement mediated search, neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal, and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, and areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. The detailed association of simultaneously obtained eye movement sequences and neural parameters is well done. These are valuable data that will contribute to our understanding of attentional modulation in visual search.

      Strengths:

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly, the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from multiple areas (V4, IT, PFC). They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, a high degree of feature and categorical specificity. This provides valuable data for the concept of a foveal-peripheral attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors, and looks towards and away from the target) and statistical rigor make these findings quite compelling.

      Weaknesses:

      While the study is generally quite strong, there are a few weaknesses to be addressed.

      (1) Little rationale is provided for recording in the selected areas, V4, IT, and PFC. Given the respective roles in sensory, object recognition, and goal-directed behavior, some rationale for this design should be offered, and commonalities/distinctions between these areas should be discussed.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction as follows:

      “V4 and inferotemporal cortex (IT), as the middle and high-level areas of the ventral visual stream, are important for object recognition and categorization [27-34], and their roles have been extensively studied in central vision. At the neuronal level, however, most investigations have largely neglected their functions during active, free-gaze visual search. The prefrontal cortex, including LPFC, has long been implicated as a source of top-down signals that bias the selection of attended features and modulate visual cortical responses [6, 9, 11, 35-40]. Although target-related visual responses have been reported in IT during visual exploration [41], and target-selective responses have been observed in the human medial temporal lobe (MTL) [42] and medial frontal cortex (MFC) [43] during visual search, these studies did not map the receptive fields (RFs) of recorded neurons.”

      We also added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area—consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      (2) Given the reliance of all analyses on saccadic behavior (towards target/distractor, towards/away from target), additional description and summaries of eye movement behavior during single trials and across trials should be provided.

      We thank the reviewer for this helpful suggestion. We have added a description of saccade behavior to the Results as follows:

      “The mean number of saccades monkeys made to find the target after the onset of the search array was 2.25 ± 1.35 (mean ± SD across trials; Table 1) of correct trials, and the mean saccade amplitude was 7.99° ± 3.58° (mean ± SD across saccades; Table 1). Monkeys could fixate on each distractor or the target freely, provided they did not maintain fixation on the target for longer than 800 ms. Across sessions, 42.44% ± 3.6% of saccades were directed to distractors, 57.56% ± 3.6% to targets, and 12.59% ± 3.46% were saccades away from targets (see our previous studies [44-46] for detailed behavioral analyses).”

      We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search task.

      We have also included Table 1, which summarizes eye movement behaviors.

      (3) The dependency of findings on top-down (categorical & feature-specific) task design should be discussed.

      We thank the reviewer for the suggestion and added a discussion as follows:

      “In this task, attention is strongly guided by top-down goals, which bias processing toward behaviorally relevant features and object categories [2, 50, 51]. Top-down attention, including categorical and feature-specific components, has been shown to modulate neural processing across the visual pathway based on task demands and to originate from distributed frontoparietal control networks [11, 35-38, 40]. Our study provides further insight into the mechanisms of goal-directed visual attention, as it is among the first to demonstrate foveal feature attention effects during free-gaze visual search, as well as the distribution of feature and spatial attention across the entire visual field.”

      Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper, including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Figure 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provides a dataset that is well-suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      We thank the reviewer for the helpful suggestion and apologize for not explicitly providing essential information about the RFs of the units. We added a detailed description of RF properties to the Results as follows:

      “The RFs of these peripheral units were further mapped using a visually guided saccade task and quantified by the number of stimuli that activated each unit (Fig. 1F-K). The eccentricities of the peripheral RFs were 6.22° ± 1.31° (mean ± SD) in V4, 7.04° ± 1.52° in IT, and 6.68° ± 1.56° in LPFC. The sizes of the peripheral RFs were 3.67° ± 1.87° in V4, 6.86° ± 3.11° in IT, and 8.65° ± 3.02° in LPFC. The numbers of items from the search array falling within peripheral RFs were 1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC (also see our previous study [44]).”

      The reviewer is correct that multiple items from the search array did fall within the RFs of peripheral-RF units. However, for focal foveal units, only the fixated stimulus fell within the RF, due to the design of the search array and the definition of these units used in our analyses (see our reply to Reviewer 1, Public Review, Question 2 for details). We agree with the reviewer that attentional modulation is typically stronger when multiple stimuli fall within RFs. In our design, peripheral RFs, on average, contained more stimuli than foveal RFs. Therefore, this difference in RF size would, if anything, be expected to bias toward stronger attentional modulation in peripheral units. This would make our observation conservative, thereby further supporting rather than undermines our main finding of robust feature-based attentional enhancement in foveal units, challenging the prevailing view that such modulation is predominantly peripheral. However, we agree that, when comparing the latency of attentional effects across brain regions in Fig. 3, we cannot rule out the influence of the number of stimuli arising from differences in RF size.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Figure 2). As a result, the reported attentional modulation coincides with the preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19), therefore, likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Figure S3) partially mitigate this concern by demonstrating that featurebased modulation persists through saccade execution.

      We thank the reviewer for raising this important question. We agree that the temporal overlap of visual, motor planning, target recognition, and behavioral relevance signals with attention can result in mixed activity, which needs to be dissociated. Therefore, when calculating feature-based attention, we did implement a series of controls. We added a discussion as follows:

      “A major challenge in interpreting neural activity related to attentional modulation is the inherent temporal overlap of visual processing, motor planning, and target recognition signals in the free-gaze visual search task [73]. To isolate genuine feature-based attention from potential confounds, we applied several stringent analytical constraints, consistent with prior studies [3, 5, 6]. Specifically, by restricting our analysis to fixations where the subsequent saccade was directed away from the RFs, we dissociated attentional modulation from the preparatory motor activity associated with saccade execution. Furthermore, by comparing responses to the same physical stimulus alternating its role as a target or distractor across trials we eliminated any potential bias introduced by stimulus identity or physical category. We restricted our analysis to fixations preceding target selection that is, before the monkeys made a behavioral choice to minimize contributions from target detection or decision-related signals.”

      We thank the reviewer for pointing out the issue of different timescales for target versus distractor fixations. To address this, we conducted a control analysis by computing foveal feature-based attentional modulation using fixations on targets and distractors with matched fixation durations. We obtained similar results. We have updated Fig. S2 to include this control analysis.

      We also clarified this point in the Results as follows:

      “We also obtained similar results when controlling for fixation durations on targets and distractors (i.e., there was no significant difference between fixation durations on targets and distractors; Wilcoxon signed-rank test, P > 0.05; Fig. S2K–P).”

      Lastly, as the reviewer correctly pointed out, the interpretation that foveal feature-based attention facilitates prolonged fixation on the target was not supported. We have revised the Results as follows:

      “On average, target fixations (256.69 ± 197.44 ms [mean ± SD]) were significantly longer than distractor fixations (156.26 ± 45.94 ms; Wilcoxon rank-sum test, P < 0.0001), and during these prolonged target fixation, foveal feature-based attention modulation was consistently observed.”

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Figures 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      We thank the reviewer for this important question. We performed a directional control analysis by computing spatial attentional modulation using paired fixations from the attention-in and attention-out conditions. Only saccades directed in nearly opposite directions—defined as having a saccade direction angle ≥ 170° within the 0–180° range—were included. We obtained similar results (Author response image 1). 

      Author response image 1.

      Peripheral spatial attentional modulation in V4, IT, and LPFC. Population response to stimuli followed by saccades directed into their RFs (attention in) versus directed approximately opposite and outside their RFs (attention out), shown for V4 (A), IT (B), and LPFC (C). Shaded area denotes ±SEM across units.

      We did control for feature-based attention when calculating spatial attentional modulation. We apologize for the lack of clarity and have added a description of this control to the Methods as follows:

      “The saccade-target stimulus in the RF during attention-in fixations was matched to a stimulus in the same location during attention-out fixations; in both conditions, this stimulus always served as a distractor for that trial, except in the “Distractor fixations to T” condition (Fig. 5 and Fig. S4), in which it instead served as the target. This design eliminates differences due to feature-based attention between the attention-in and attention-out conditions.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3C: Unclear how to compare LPFC vs V4 for foveal units since only data from peripheral LPFC is shown?

      We thank the reviewer for pointing out this mistake. In Fig. 3C, we only compared LPFC peripheral units, V4 peripheral units, and V4 foveal units. We have corrected this in the legend of Fig. 3 as follows:

      “Shown are cumulative distributions of feature-attention effect latencies, computed from individual foveal face-, house-, and non-selective units in V4 and IT, and from peripheral non-selective units in V4,

      IT, and LPFC.”

      (2) On page 8, last para: For units with peripheral RFs ... Is this controlled for whether the saccade is to targets or to distractors?

      We thank the reviewer for the question. We indeed addressed this concern by separating fixations based on whether the subsequent saccade was directed to a target or a distractor, and by analyzing attention modulation within each condition. Therefore, attention effects were evaluated while holding the saccade destination constant, effectively controlling for potential confounds related to saccade target selection.

      (3) Page 9: The authors find that target fixations were longer than distractor fixations and conclude that this supports the idea that foveal feature-based attention increases fixation duration, but this interpretation is pure conjecture, and there is no experimental manipulation presented in this paper that helps to establish this interpretation.

      We thank the reviewer for this important comment. We agree that this observation does not, by itself, support our original interpretation, and we have modified it in the Results. Please refer to the last paragraph of our Reply to Question 2 from Reviewer 3 (Public Review).

      (4) Data analysis: receptive field. The authors state that visual response to a cue and the stimulus array was assessed during the 0-200 ms window after stimulus onset. However, after the array onset, the animal could saccade within the 200 ms window. How do the authors ensure uniform stimulation during the 0-200 ms window?

      We thank the reviewer for this question. The activity of units in V4, IT, and LPFC within the 200 ms window after array onset primarily reflected visual stimulation prior to saccades, because typical saccade latencies were approximately 150–200 ms, and the response onset latencies of these units were around 50 ms.

      (5) On page 19, the authors state that to assess feature attention in peripheral RFs, they divided trials into target and distractor fixations. In the former, there was a target in the neuron's RF. This is confusing. I assume target fixations imply fixating on a target, but the authors may mean fixations where a target is in the RF. Please clarify.

      We thank the reviewer for pointing out this confusion. In the original manuscript, we intended to sort fixations by whether a target stimulus was located within the unit’s peripheral RF. To avoid further confusion, we have revised the description in the Methods as follows:

      “we sorted fixations during the search period, following a procedure similar to that in our previous study [5], into two types: “target” – a target stimulus was located within the unit’s peripheral RF; and “distractor” – the same stimulus appeared in the same peripheral RF location but served as a distractor.”

      (6) Figure S1: Are these example units? How many trials? SEM? The sharp rise and no noise are inconsistent; the former suggests minimal smoothing, while the latter suggests lots of smoothing.

      We thank the reviewer for these questions. We showed average responses across all units in Fig. S1. On average, there were 941.79 ± 182.56 trials (mean ± SD across sessions). Shaded areas indicate ±SEM across units. The sharp rise reflects the synchronous response of neurons to the stimulus, while the smooth appearance and low noise result from averaging across a very large number of units and trials.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) One weakness of this manuscript is the lack of a rationale for choosing V4, IT, and PFC. Specifically, what are the predictions of the roles of these respective areas in the integration of current and peripheral (future foveal) views? There is a significant literature linking the pre-saccadic peripheral stimulus and the post-saccadic foveal stimulus, suggesting that both spatial and temporal integration occur. However, whether such integration occurs at high or low cortical levels is unknown. By recording from mid-tier (V4) and high-order areas (IT, PFC), the authors have an opportunity to address this question. However, there is no mention of this topic, either in the introduction, results, or discussion. I find this omission surprising. At the very least, it should contribute to experimental design rationale and some discussion.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction and a discussion about this integration. Please refer to our Reply to Question 1 from Reviewer 2 (Public Review).

      (2) As both behavior and neural recordings are collected, a figure on saccadic patterns would enhance the reader's understanding. Questions that come to mind are: What does a single search trial look like? How many saccades are there per trial? How often is the target identified after 1, 2, 3, etc saccades? What is the average size of a saccade? Although this is not a study of search strategy per se, a modicum of description of the search sequences would provide context on the behavior. I suggest an illustration of one or more sample trials; a summary of saccade behavior would also be helpful for understanding the data in relation to behavioral performance.

      We thank the reviewer for this helpful suggestion. We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search, providing an example of a single search trial. Additionally, we have added a description of saccade behavior to the Results and included Table 1, which summarizes eye movement behavior. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for further details.

      (3) "Consistently, the probability of making a saccade to a peripheral target was higher following distractor fixations (75.22%) than following target fixations (48.44%, or 63.49% after probability calibration; see Methods), indicating the important role of peripheral feature-based attention in guiding eye movements" It should be noted that this target-oriented visual search is fundamentally a top down task. Once the target is found, the reward is obtained; saccades to distractors are not rewarded, so saccades are more likely. So certainly this task design would increase the post-distractor saccades and decrease the number of post-target saccades. Please clarify the behavioral paradigm: once a reward is obtained, does the task continue, or is a new trial initiated?

      We apologize for the confusion regarding the behavioral paradigm. We would like to clarify that when the target was found and fixated for 800 ms, the reward was delivered and no further saccades occurred. However, if the target was not fixated for 800 ms, the search could continue. It is worth noting that the target fixations in our analyses were restricted to those occurring during ongoing search behavior, excluding target fixations associated with trial termination and reward delivery. Moreover, we compared the probability of making a saccade to the target, rather than the absolute number of saccades, following these fixations. We have modified the Results for clarification, as follows:

      “Two monkeys performed a category-based visual search task, where their objective was to fixate on one of the two search targets that matched the category of the cue (Fig. 1A, B). Specifically, the monkeys were presented with a central fixation point for 400 ms, followed by a cue lasting 500-1300 ms. After a 500 ms delay, a search array appeared with 11 items, including two targets, randomly chosen from 20 possible locations (Fig. 1E). The monkeys had 4000 ms to find one target and maintain fixation on it for 800 ms to earn a juice reward. Fixating on either target completed the trial, and the monkeys did not search for the second target. A new trial began after the reward. It is worth noting that the two target stimuli matched the category of the cue but were different images. The monkeys were required to maintain fixation throughout the cue and delay periods. During search, however, eye movements were unconstrained, and monkeys could revisit each search distractor or target as long as they did not fixate on a target for 800 ms.”

      (4) The fact that there are many more peripheral units in LPFC suggests that this is a region of foveal/periph integration. Combined with the finding that the LPFC leads the attentional effects, this should be a discussion point.

      We thank the reviewer for the suggestion and we added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      Minor comments:

      (1) Figures 2A-D. "These face-selective units also showed slightly enhanced responses to house targets in IT (P < 0.05), but not in V4 (P = 0.89)." It does not appear enhanced.

      We agree with the reviewer that the effect is modest and does not appear strongly enhanced. However, the average response in the 150–225 ms time window to the house target was significantly higher than that to the house distractor in IT face-selective units (Wilcoxon signed-rank test, P = 0.042). We modified the description in the Results as follows: 

      “These face-selective units also showed weakly but significantly enhanced responses to house targets in IT (P < 0.05)”

      (2) Figure 3. For population comparison, a bootstrapped null distribution was used, and a 2-sided permutation test was used to determine the latency difference between the target and distractor; please show these results (described in text) in a figure. Figures 3A-C are described as the latency of individual units. So each of these graphs is the mean of multiple units? So this is also a population analysis? What is the difference between these two comparisons? This is somewhat confusing.

      We apologize for the confusion and thank the reviewer for pointing this out. Each panel in Fig. 3 shows the cumulative distribution of latencies across individual units within each brain region, reflecting the variability of response timing across single neurons. For this analysis, we first calculate the latency of each unit separately. In contrast, population-level latency is measured from the averaged responses of all units within each region (Fig. 2), which captures the overall timing of the population response rather than individual variability. Statistical comparisons at the population level are performed using a two-sided permutation test. We modified Fig. 2 to better illustrate the population-level latency results.

      (3) Did peripheral RFs span more than a single stimulus in the array? If so, how does this impact the interpretation of Figure 5?

      We thank the reviewer for pointing this out. The reviewer is correct that, in peripheral RFs, more than one stimulus from the search array could fall within the receptive field (1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC). We controlled for this in our analysis of both feature-based and spatial attention effects for peripheral units in Fig. 5. For feature-based attention, we performed the analysis in a stimulus-by-stimulus manner within each category (house and face), such that when a given stimulus served as the target, it was the only target within the RF, and when it served as a distractor, it was the only distractor of its category within the RF. Although additional distractor could still fall within the RF, their identities were random across conditions and thus would be averaged out. A similar approach was applied to spatial attention, where the stimulus-by-stimulus comparison was extended across all four categories, and attention-out stimuli were paired with the corresponding saccade-target stimuli in the attention-in condition, with the effects of other randomly present distractors averaged out. Therefore, the effects shown in Fig. 5 reflect comparisons at the level of individual stimulus, minimizing confounds from other stimuli within the RF.

      (4) Figure 5G: "during "Target fixations to D", there was no significant feature attentional enhancement in response to the peripheral target (Wilcoxon signed-rank test, P > 0.05; Figure 5G-I left panels). It appears that there is some effect of spatial attention during Target Fix to D trials.

      We thank the reviewer for pointing this out and have revised the Results as follows:

      “We further found that spatial attentional enhancements to the saccade target were reduced during target fixations compared to distractor fixations in V4 and IT when activity was aligned to fixation onset (Wilcoxon rank-sum test, P < 0.05; Fig. 5G, H versus Fig. 5A, B), although this effect was not completely abolished.”

      (5) The specific areas of IT and LPFC that were recorded should, as much as possible, be mentioned.

      We thank the reviewer for the helpful suggestions and have added a description of the specific IT and LPFC recording sites to the Methods as follows:

      “Recordings in IT spanned the central IT cortex, encompassing the area between the anterior middle temporal sulcus (AMTS) and the posterior middle temporal sulcus (PMTS), including TE and TEO. Recordings in LPFC were located anterior to the arcuate sulcus (AS) and lateral to the principal sulcus (PS), mainly covering areas 45 and 44.”

      (6) It is often difficult to distinguish the different lines, e.g., red solid vs red dotted, due to their overlap. Would the removal of the error band make this clearer? If so, could put full figure with error bands in the Supplementary Figure.

      We thank the reviewer for this helpful suggestion. To improve visual clarity, we adjusted Fig. 6, Fig. 7, Fig. S2, Fig. S3, Fig. S4, and Fig. S6 by changing the line styles and placing the shaded error bands beneath the traces, allowing the lines to remain clearly visible despite overlap.

      (7) For easy access, the number of saccades to/from targets/distractors should be put into a table.

      We thank the reviewer for the suggestion. We calculated the probability of saccades to and from targets and distractors for each session and report the mean ± SD across sessions in Table 1, as the mean number of saccades per trial was only 2.3. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for Table 1.

      Reference

      (1) O'Craven, K.M., P.E. Downing, and N. Kanwisher, fMRI evidence for objects as the units of attentional selection. Nature, 1999. 401(6753): p. 584-7.

      (2) Baldauf, D. and R. Desimone, Neural mechanisms of object-based attention. Science, 2014. 344(6182): p. 424-7.

      (3) Hayden, B.Y. and J.L. Gallant, Combined effects of spatial and feature-based attention on responses of V4 neurons. Vision Res, 2009. 49(10): p. 1182-7.

      (4) Bichot, N.P., et al., A Source for Feature-Based Attention in the Prefrontal Cortex. Neuron, 2015. 88(4): p. 832-844.

      (5) Reddy, L. and N. Kanwisher, Category selectivity in the ventral visual pathway confers robustness to clutter and diverted attention. Curr Biol, 2007. 17(23): p. 2067-72.

      (6) Peelen, M.V., L. Fei-Fei, and S. Kastner, Neural mechanisms of rapid natural scene categorization in human visual cortex. Nature, 2009. 460(7251): p. 94-7.

      (7) Cukur, T., et al., Attention during natural vision warps semantic representation across the human brain. Nat Neurosci, 2013. 16(6): p. 763-70.

      (8) Keller, A.S., et al., Attention enhances category representations across the brain with strengthened residual correlations to ventral temporal cortex. Neuroimage, 2022. 249: p. 118900.

      (9) Zhang, J., et al., Behavioral and Neural Mechanisms of Face-Specific Attention during GoalDirected Visual Search. The Journal of Neuroscience, 2024. 44(46): p. e1299242024.

      (10) Bichot, N.P., A.F. Rossi, and R. Desimone, Parallel and serial neural mechanisms for visual search in macaque area V4. Science, 2005. 308(5721): p. 529-534.

      (11) Bichot, N.P., et al., The role of prefrontal cortex in the control of feature attention in area V4. Nat Commun, 2019. 10(1): p. 5727.

      (10) Cohen, M.R. and J.H. Maunsell, Using neuronal populations to study the mechanisms underlying spatial and feature attention. Neuron, 2011. 70(6): p. 1192-204.

      (11) Maunsell, J.H. and S. Treue, Feature-based attention in visual cortex. Trends Neurosci, 2006. 29(6): p. 317-22.

      (12) McAdams, C.J. and J.H. Maunsell, Attention to both space and feature modulates neuronal responses in macaque area V4. J Neurophysiol, 2000. 83(3): p. 1751-5.

      (13) Motter, B.C., Saccadic momentum and attentive control in V4 neurons during visual search. J Vis, 2018. 18(11): p. 16.

      (14) Sapountzis, P., S. Paneri, and G.G. Gregoriou, Distinct roles of prefrontal and parietal areas in the encoding of attentional priority. Proc Natl Acad Sci U S A, 2018. 115(37): p. E8755-E8764.

      (15) Treue, S. and J.C. Martinez Trujillo, Feature-based attention influences motion processing gain in macaque visual cortex. Nature, 1999. 399(6736): p. 575-9.

      (16) Zhou, H. and R. Desimone, Feature-based attention in the frontal eye field and area V4 during visual search. Neuron, 2011. 70(6): p. 1205-17.

    1. eLife Assessment

      One-carbon tetrahydrofolate metabolism plays a crucial role in producing essential metabolic intermediates. In this valuable study, the authors employ a solid genetics-based approach to demonstrate that three distinct metabolic pathways are essential for synthesising 1C-tetrahydrofolates (1C-THF). Disrupting any of these pathways impairs both growth and virulence.

    2. Reviewer #1 (Public review):

      Summary:

      This study identifies three redundant pathways-glycine cleavage system (GCS), serine hydroxymethyltransferase (GlyA), and formate-tetrahydrofolate ligase/FolD-that feed the one-carbon tetrahydrofolate (1C-THF) pool essential for Listeria monocytogenes growth and virulence. Reactivation of the normally inactive fhs gene rescues 1C-THF deficiency, revealing metabolic plasticity and vulnerability for potential antimicrobial targeting.

      Strengths:

      (1) Novel evolutionary insight-Reversible reactivation of a pseudogene (fhs) shows adaptive metabolic plasticity, relevant for pathogen evolution.

      (2) They systematically combine targeted gene deletions with suppressor screening to dissect the folate/one-carbon network (GCS, GlyA, Fhs/FolD).

    3. Reviewer #3 (Public review):

      Summary:

      In this study, Freier et al., demonstrate that 3 distinct metabolic pathways are critical for the synthesis of 1C-THF, a metabolite that is crucial for the growth and virulence of Listeria monocytogenes. Using an elegant suppressor screen, they also demonstrate the hierarchical importance of these metabolic pathways with respect to the biosynthesis of 1C-THF.

      Strengths:

      This study uses elegant bacterial genetics to confirm that 3 distinct metabolic pathways are critical for 1C-THF synthesis in L. monocytogenes and lack of either one of these pathways compromises bacterial growth and virulence. The study uses a combination of in vitro growth assays, macrophage-CFU assays and murine infection models to demonstrate this.

      Comments on revisions:

      The revised manuscript is improved, and the additional genetic experiments provide further support for the proposed metabolic model. However, the central conclusion is not fully established without direct measurement of 1C-THF levels. While I appreciate the authors' explanation regarding the technical limitations, quantitative metabolite measurements (e.g., by mass spectrometry) would have provided much stronger evidence linking the genetic perturbations to altered 1C-THF pools.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study identifies three redundant pathways-glycine cleavage system (GCS), serine hydroxymethyltransferase (GlyA), and formate-tetrahydrofolate ligase/FolD-that feed the one-carbon tetrahydrofolate (1C-THF) pool essential for Listeria monocytogenes growth and virulence. Reactivation of the normally inactive fhs gene rescues 1C-THF deficiency, revealing metabolic plasticity and vulnerability for potential antimicrobial targeting

      Strengths:

      (1) Novel evolutionary insight - reversible reactivation of a pseudogene (fhs) shows adaptive metabolic plasticity, relevant for pathogen evolution.

      (2) They systematically combine targeted gene deletions with suppressor screening to dissect the folate/one-carbon network (GCS, GlyA, Fhs/FolD).

      Weaknesses:

      (1) The study infers 1C-THF depletion mostly genetically and indirectly (growth rescue with adenine) without direct quantification of folate intermediates or fluxes. Biochemical confirmation, LC-MS-based metabolomics of folates/1C donors, or isotopic tracing would strengthen mechanistic claims.

      We agree with the reviewer that quantification of 1C-THF intermediates would strengthen our conclusions. However, the chemical methodologies to extract folates from L. monocytogenes are not established in our lab. Moreover, quantification of C1-substituted folates requires comprehensive biochemical and analytical expertise that we also do not have and which we cannot cover though co-operations. However, to further strengthen our arguments, we have introduced an experiment in the updated manuscript that demonstrates synthetic lethality of a ΔgcvPAB ΔglyA mutant with a deletion of fold (Fig. 7A). This gene encodes 5,10-methylene-tetrahydrofolate dehydrogenase/ 5,10-methylene-tetrahydrofolate cyclohydrolase, which is the third enzyme involved in N5, N10-methylene-THF generation next to GlyA and GcvPBA. Synthetic lethality of a ΔgcvPAB ΔglyA double mutant with a fold deletion is best explained by GcvPAB and GlyA also feeding the N5,N10-methylene-THF pool.

      (2) In multiple result sections, the authors report data from technical triplicates but do not mention independent biological replicates (e.g., Figure 2C, Figure 4A-B, Figure 6D). In addition, some results mention statistical significance but without a detailed description of the specific statistical tests used or replicates, such as Figure 2A-C, Figure 2E, and Figure 2G-I.

      We thank the reviewer for this helpful comment. Experiments were usually repeated three independent times, with each repetition including three technical replicates. Mean values and standard deviations were usually calculated from the technical replicates of a representative run. Statistical significance was calculated using t-tests for pairwise comparisons or t-tests using Bonferroni-Holm correction for multiple comparisons. We made sure that this is explicitly explained for each experiment in the figure legends. Wherever other calculations were used, we also clarified this in the legends.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Freier et al examines the impact of deletion of the glycine cleavage system (GCS) GcvPAB enzyme complex in the facultative intracellular bacterial pathogen Listeria monocytogenes. GcvPAB mediates the oxidative decarboxylation of glycine as a first step in a pathway that leads to the generation of N5, N10-methylene-Tetrahydrofolate (THF) to replenish the 1-carbon THF (1C-THF) pool. 1C-THF species are important for the biosynthesis of purines and pyrimidines as well as for the formation of serine, methionine, and N-formylmethionine, and the authors have previously demonstrated that gcvPAB is important for bacterial replication within macrophages. A significant defect for growth is observed for the gcvPAB deletion mutant in defined media, and this growth defect appears to stem from the sensitivity of the mutant strain to excess glycine, which is hypothesized to further deplete the 1C-THF pool. Selection of suppressor mutations that restored growth of gcvPAB deletion mutants in synthetic media with high glycine yielded mutants that reversed stop codon inactivation of the formatetetrahydrofolate ligase (fhs) gene, supporting the premise that generation of N10-formyl-THF can restore growth. Mutations within the folk, codY, and glyA genes, encoding serine hydroxymethyltransferase, were also identified, although the functional impact of these mutations is somewhat less clear. Overall, the authors report that their work identifies three pathways that feed the 1C-THF pool to support the growth and virulence of L. monocytogenes and that this work represents the first example of the spontaneous reactivation of a L. monocytogenes gene that is inactivated by a premature stop codon.

      Strengths:

      This is an interesting study that takes advantage of a naturally existing fhs mutant Listeria strain to reveal the contributions of different pathways leading to 1C-THF synthesis. The defects observed for the gcvPAB mutant in terms of intracellular growth and virulence are somewhat subtle, indicating that bacteria must be able to access host sources (such as adenine?) to compensate for the loss of purine and fMet synthesis. Overall, the authors do a nice job of assessing the importance of the pathways identified for 1C-THF synthesis.

      Weaknesses:

      (1) Line 114 and Figure 1: The authors indicate that the gcvPAB deletion forms significantly fewer plaques in addition to forming smaller plaques (although this is a bit hard to see in the plaque images). A reduction in the overall number of plaques sounds like a bacterial invasion defect - has this been carefully assessed? The smaller plaque size makes sense with reduced bacterial replication, but I'm not sure I understand the reduction in plaque number.

      The observation that the ΔgcvPAB mutant forms fewer plaques was not our claim, and we have already addressed the possibility of an invasion defect by quantifying bacterial numbers during infection of 3T3 cells. As shown in Fig. 2A, the ΔgcvPAB mutant invades 3T3 cells (the same cells used in the plaque formation assays) as efficiently as the wild type but exhibits reduced intracellular growth. Therefore, the plaquing defect is not due to impaired invasion. Furthermore, we also have analyzed the intracellular dissemination of the ΔgcvPAB mutant in 3T3 fibroblasts compared to a ΔactA mutant by microscopy. This shows that the ΔgcvPAB is evenly distributed throughout the infected host cells as the wild type and unlike the ΔactA mutant, which forms clusters (Fig. S1). Both experiments indicate that the reduced plaque area results from impaired intracellular growth rather than defects in invasion or cell-to-cell spread. The apparent reduction in plaque numbers in ΔgcvPAB-infected 3T3 cells is likely due to a strong reduction in plaque size, with only the largest plaques remaining visible. We have rephrased the relevant section to clarify this point and avoid any potential confusion:

      “In agreement with our previous results, only small plaques were formed in 3T3 cells upon infection with the ΔgcvPAB mutant (plaque area: 15±19% of wild-type level) and small plaques were also formed by the complemented strain in the absence of IPTG (50±19%).”

      (2) Do other Listeria strains contain the stop codon in fhs? How common is this mutation? That would be interesting to know.

      We determined the frequency of inactivated fhs genes among 30,000 publicly available L. monocytogenes genomes. The analysis identified only 10 isolates carrying truncated fhs alleles. These isolates fell into two groups: (i) EGD-e and its descendants, and (ii) a cluster of five ST2 food isolates. These findings have been added as a separate results section.

      (3) Based on the observation that fhs+ ΔgcvPAB ΔglyA mutant is only possible to isolate in complex media, and fhs is responsible for converting formate to 1C-THF with the addition of FolD, have the authors thought of supplementing synthetic media with formate and assessing mutant growth?

      No, we did not test formate supplementation. However, we included additional experiments testing adenine and thymine supplementation (Fig. 6E and 7B). These results show that purine and thymine become limiting in mutants lacking 1C-THF synthesizing pathways.

      Reviewer #3 (Public review):

      Summary:

      In this study, Freier et al. demonstrate that 3 distinct metabolic pathways are critical for the synthesis of 1C-THF, a metabolite that is crucial for the growth and virulence of Listeria monocytogenes. Using an elegant suppressor screen, they also demonstrate the hierarchical importance of these metabolic pathways with respect to the biosynthesis of 1C-THF.

      Strengths:

      This study uses elegant bacterial genetics to confirm that 3 distinct metabolic pathways are critical for 1CTHF synthesis in L. monocytogenes, and the lack of either one of these pathways compromises bacterial growth and virulence. The study uses a combination of in vitro growth assays, macrophage-CFU assays, and murine infection models to demonstrate this.

      Weaknesses:

      (1) The primary finding of the study is that the perturbation of any of the 3 metabolic pathways important for the synthesis of 1C-THF results in reduced growth and virulence of L. monocytogenes. However, there is no evidence demonstrating the levels of 1C-THF in the various knockouts and suppressor mutants used in this study. It is important to measure the levels of this metabolite (ideally using mass spectrometry) in the various knockouts and suppressor mutants, to provide strong causality.

      As already outlined above, we do not have any experimental possibilities to measure 1C-substituted THF in L. monocytogenes extracts directly. However, to provide additional evidence for our interpretation that “Three pathways feed the 1C-THF pool…”, we included additional genetic experiments.

      The first experiment demonstrates that the growth defect of the fhs- ΔglyA igcvPAB strain in synthetic medium lacking IPTG—which reflects the synthetic lethality of fhs with glyA and gcvPAB—can be rescued by the addition of adenine (Fig. 6E). This indicates that the fhs/fold pathway, GlyA, and the glycine cleavage system are essential due to their combined contribution to purine biosynthesis. This result confirms the metabolic model presented in Fig. 1A and thus supports the hypothesis that all three pathways contribute to 1C-THF biosynthesis.

      The second experiment additionally demonstrates synthetic lethality of gcvPAB and glyA with the fold gene. FolD acts downstream of Fhs and is one of the three enzymes synthesizing N5, N10methylene-THF shown in Fig. 1A. The fold gene is essential in EGD-e (PMID: 36114002), most likely explained by fhs inactivation. However, we were able to delete fold in an EGD-e background carrying a reconstituted fhs gene and the resulting fhs<sup>+</sup> Δfold strain was as viable as a fhs<sup>+</sup> ΔglyA ΔgcvPAB strain (Fig. 7A). However, a fhs<sup>+</sup> Δfold ΔglyA igcvPAB strain required IPTG for growth in BHI medium (Fig. 7A), indicating that the simultaneous deletion of glyA and gcvPAB is not tolerated in the absence of fold, similar to what is observed in the absence of fhs. Notably, this growth defect was not rescued by adenine supplementation (Fig. 7B), but was alleviated by thymine addition, which is also consistent with the metabolic model shown in Fig. 1A.

      Even though we are unable to directly demonstrate reduced 1C-THF levels, we hope that these two genetic approaches together with the revised title and heading of the relevant paragraph will, in the reviewers’ eyes, support our hypotheses.

      (2) The story becomes a little hard to follow since macrophage-CFU assays and murine infection model data precede the in vitro growth assays. The manuscript would benefit from a reorganization of Figures 2,3, and 4 for better readability and flow of data.

      We respectfully disagree with the reviewer. The attenuation of the ΔgcvPAB mutant in macrophages and fibroblasts was the primary motivation for further investigating its phenotype. Therefore, we chose to begin the manuscript with results from various virulence studies before presenting the sections that provide mechanistic explanations. In our view, this sequence represents a more logical and coherent way to present the data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Synthetic medium assumptions: LSM "mimics" intracellular limitation but isn't chemically validated against host cytosolic composition. Nutrient availability conclusions could be biased.

      This is correct. We added this information:

      “LSM broth is a chemically defined medium that contains all components required for growth at defined concentrations, but it has not been chemically validated to reflect host cytosolic conditions…”

      (2) Glycine toxicity: The paper introduces the concept of glycine toxicity in ΔgcvPAB mutants. However, the conditions under which glycine becomes toxic could be further elucidated. Why glycine causes toxicity despite being an essential metabolite in other contexts requires a more in-depth mechanistic explanation.

      The concept of glycine toxicity in GCS mutants has been described previously by other researchers. We have added further details to better explain glycine toxicity and how it accounts for the growth phenotype of the ΔgcvPAB mutant:

      “If glycine cannot be catabolized (and 1C-THF cannot be generated) by the GCS due to deletion of gcvPAB, glycine might be re-routed to the serine hydroxymethyl transferase GlyA for serine formation, even though this would consume 1C-THF and therefore even further deplete the cell for 1C-THF.” and

      “In the complete absence of glycine, growth of the ΔgcvPAB mutant was largely unaffected, presumably because glycine cannot be converted to serine by GlyA anymore, thereby conserving the 1C-THF pool.”

      (3) Figures:

      (a) Figure 1B, scale bar?

      A scale bar was added.

      (b) Figure 1C, the standard error bars for igcvPAB (both with and without IPTG) are relatively wide, indicating high variability in the data. This suggests that the results in the igcvPAB group are not as consistent as the wild-type (wt) or ΔgcvPAB groups. Please show the original data and perform a statistical test (e.g., t-test or ANOVA).

      The original data have been added to Fig. 1B. t-test results (with Bonferroni-Holm correction) are now included for all samples.

      (c) Figure 2D, scale bar?

      These are sections of agar plates. From our point of view, a scale bar does not add relevant information.

      (d) Figure 3, please check the labels of the Figure 3 legend. (D) and (E) or A-B?

      Thanks, corrected.

      (e) Figure 5E, quantification of plaque areas?

      The plaque areas were quantified. A blot showing these quantitative data is now presented in Fig. 5F.

      Reviewer #2 (Recommendations for the authors):

      Line 58: There are published studies that indicate that syncytiotrophoblasts are actually resistant to Listeria infection and that it is extravillous trophoblasts that are likely to serve as entry points for Listeria into the placenta (see, for example, Lowe et al, Infect Immun. 2018 Volume 86 Issue 6 e00801-17).

      We thank the reviewer for this comment and have removed our statement claiming that syncytiothrophoblasts are the entry point as this is not relevant to the understanding of the work presented here.

      Reviewer #3 (Recommendations for the authors):

      (1) Please mention the number of times experiments were performed as independent biological replicates, wherever applicable.

      We added this information to the figure legends wherever this was necessary.

      (2) Please provide the details of the type of statistical analysis used for the various graphs, either in the figure legends or in the materials & methods section.

      This information was also added to the figure legends wherever it still was missing.

      (3) Can the authors comment on how the weight-loss phenotype of animals and the variation in the size of the spleen between animals infected with wild-type and mutants in Figure 2 can be explained without any significant changes in the CFU? Additionally, I did not see details regarding the number of animals used in the murine infection model experiments. Please mention this along with the type of statistical analyses used.

      We do not see a contradiction here, as the apparent differences are explained by the distinct time points at which CFU numbers (day 3 post-infection) and spleen size (day 9 postinfection) were measured. Starting from day 8, the difference in weight between animals infected with the wild-type strain and those infected with the ΔgcvPAB mutant becomes clear for the first time. At day 3, no significant differences are detected in either CFU numbers or weight. By day 9, when the weight difference has become apparent, differences in spleen size are also observed. To improve clarity, the time points at which each analysis was performed have been added to Fig. 2G and Fig. 2I. The number of infected animals and the type of statistical analysis used are now specified in the figure legend.

    1. eLife Assessment

      This is a detailed and well-designed simulation study of the utility of replication metrics in animal-to-human study translations in bridging the gap between laboratory discoveries and health practice, a critical consideration in turning laboratory scientific research findings into tangible, real-world applications, to directly help human health. The study approaches are convincing, and the findings are important, as they offer insights into clinical research translations to advance health decision-making.

    2. Reviewer #1 (Public review):

      [Editors' note: This revised version of your article has been assessed by the Reviewing Editor without further input from the original reviewers. The comments raised by the original reviewers in the earlier round of review have been addressed. The study findings are quite insightful and important, and the evidence is strong, convincing, and a substantial addition to the evidence base.]

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

    3. Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodated for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

    4. Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

      Weaknesses:

      (1) The conceptual distinction between replication and translation could be more clearly emphasized.

      (2) Interpretation of results is dense and can be challenging to follow without a clear and summarized.

      (3) Some simulation parameters (effect sizes, heterogeneity, and number of animal studies) require more substantial justification.

      (4) Practical recommendations could be more explicit to guide applied researchers.

      We thank Reviewer 1 for the general positive assessment of our study and for the constructive feedback. We have addressed all of the four identified weaknesses in the revised manuscript. Specifically,

      (1) Conceptual distinction between replication and translation. We have reinforced this distinction at multiple points in the manuscript: in Section 2.7 (just before introducing the translation success metrics), in the Discussion, and in a new working definition of translation success added to the Introduction. We further explicitly acknowledge that statistical translation success, as defined here, is narrower than biological translation.

      (2) The dense result section. We have added a summary Table (Table 3) at the end of the Results section that compares all metrics on key properties (strengths and weaknesses, overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch). We also direct readers to this table early in Section 3.2, so that readers less interested in the technical details can obtain the key take-home messages without reading the full section.

      (3) Further justification of simulation parameters. We have substantially extended the rationale for our parameter choices in Section 2.4 and the Limitations section. We explain that our parameters are grounded in an empirical meta-analytic dataset, contextualise the large effect size and heterogeneity value against published benchmarks from preclinical research, and clarify that our main goal was to explore directional trends rather than absolute performance under specific values. We have also added an invitation for others to explore alternative parameter spaces using our openly available code.

      (4) Practical recommendations. We have extended the Recommendations section (pages 21–22) with more explicit scenario-specific guidance, supported by the new summary table.

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodate for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

      We thank Reviewer 2 for the thoughtful summary and for recognising the strengths of our study. We believe that both weaknesses were addressed in the revised version of our manuscript. Specifically,

      (1) Heterogeneity in animal studies. We agree that high heterogeneity in animal studies is an important limitation, and we address it directly in our simulation design by including a wide range of heterogeneity values (including very high levels, as observed in animal studies). Our results show clearly how heterogeneity affects the performance of each metric, and we highlight this in both the new summary Table (Table 3) and the Recommendations section which was extended. We also caution applied researchers to pay special attention to heterogeneity when interpreting translation results.

      (2) Focus on the statistical component of translation. We fully agree that statistical translation success is only one aspect of a broader decision-making process. We have elaborated on this in the revised manuscript, both in a new working definition of translation success in the Introduction (which explicitly distinguishes statistical from biological translation) and in a new paragraph in the Discussion section where we situate our metrics within translational decision-making frameworks such as PATH. They make it clear that progression to human trials depends on a suite of evidence of which statistical translation is only one part.

      Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Weaknesses:

      The weaknesses of the study are the limited focus on specific metrics, the assumptions, particularly in the limited number of human study variables, and the less-than-ideal approachable summary of findings for a non-technical audience.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

      We thank Reviewer 3 for the positive assessment and for the constructive suggestions.

      We have addressed the identified weaknesses as follows:

      (1) The assumptions around human study variables. We acknowledge these as inherent constraints of the simulation design. We have added a note in the Limitations section about the fixed human sample size (N = 107 per group), clarifying that while this value is grounded in a power analysis as per regulatory standards, it represents one particular scenario and may not generalise to all contexts. Further, we have contextualised and motivated the other simulation parameters better. We also invite readers to explore alternative conditions using our openly available code.

      (2) Approachability of the summary of findings for a non-technical audience. We have added a summary Table (Table 3) at the end of the Results section, comparing the metrics on key properties including overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch. We direct readers to this table early in Section 3.2 so that those less interested in the technical details can obtain the main take-home messages without reading the full section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) Conceptual framing: clearer distinction between replication vs translation

      The Introduction correctly points out the conceptual difference between replication and translation (animal to human), but this distinction needs to be reinforced repeatedly, especially when interpreting metric performance. For instance, several metrics (e.g., meta-analysis, replication BF) inherently assume exchangeability of findings, which is rarely justified in translation because species differ biologically.

      The manuscript should explicitly state why treating animal findings as "original studies" and human findings as "replications" can be misleading. Add a subsection in the Discussion: Why replication metrics behave differently in translation settings. This will help provide a more straightforward interpretation of the results beyond the numerical findings.

      Thank you for your feedback. While the purpose of our study is to assess the applicability of the replication success metrics in the translation context, we agree that the reader should be reminded that these two concepts differ and metrics’ assumptions might not always hold. We have reiterated the difference between replication and translation in Section 2.7, just before we introduce the translation success metrics (see end of page 7). We also reiterate it in the Discussion (see end of page 20). Here, we emphasise that while some of the metrics assume that both studies investigate the same effect, this is unlikely to be the case in translation, leading to some of the metric’s assumptions being violated which impacts the performance of the metrics.

      (2) Stronger justification of the simulation parameters is needed

      The simulation factors are comprehensively presented (Table 1), but certain choices appear arbitrary or oversimplified.

      - Effect sizes: The three levels (0, −4.44, −24.37) are derived from the motivating dataset, but the paper should explain that these represent extremely large effects in many biomedical contexts.

      - Heterogeneity values: τ<sup>2</sup> = 291.1 is enormous; adding context about real-world heterogeneity distributions would help.

      - Number of animal studies (k): Using only 2-5 studies may not reflect reality; many preclinical fields have >30 studies before clinical translation.

      These choices should be more explicitly defended in Section 4 (Limitations), beyond the brief mention already there. Provide a sensitivity analysis, or explain why extrapolation beyond this parameter space is reasonable.

      We agree that the choice of the parameter values might sometimes appear arbitrary. However, instead of arbitrarily choosing parameter values, we base our choice on data from a meta-analysis. This particular meta-analysis might not be representative of all of pre-clinical and clinical research, but because Terstappen included both animal and human studies investigating the same research question it was particularly well suited. They further used an outcome (maternal blood pressure) that is comparable between rats and humans, which is quite rare. We have specified this further in Section 2.4 (Motivating dataset, page 5). In the Limitations section, we acknowledge any possibly unrealistic simulation conditions again, and emphasize that our main goal was to explore trends in the metrics’ behavior as the conditions changed rather than their absolute performance under specific values. Further, the effect sizes (0, −4.44, −24.37 mmHg) span a meaningful range on the unstandardized mean difference scale for blood pressure measurements: from no effect to a modest but clinically relevant reduction to a large effect typical of animal studies. The large heterogeneity value corresponds to a relative heterogeneity of I^2 of 95.33% in the animal meta-analysis. While this appears high, it is frequently observed in preclinical research: Hooijmans et al (2022) showed that 55% of animal study meta-analyses using mean differences as effect size measure have I^2>75%. We also added a footnote reiterating the fact that such high effect sizes (on the raw mean difference scale) are indeed common in animal studies (on page 6). Regarding k, we acknowledge that pooling only 2 to 5 animal studies may not reflect common practice. However, the directional trends in type 1 error and power are clearly visible in our Figures. Larger k decreases the type 1 error of the animal studies, while the power is increased unless there is high heterogeneity between animal studies and there is only a small effect. Extending the range further is unlikely to change the conclusions. Moreover, in practice, the decision to advance to human trials considers evidence well beyond the statistical considerations we simulate. All of the above is now emphasized more explicitly in both the methods, where we have substantially extended the reasoning for choosing the simulation conditions, and the limitations section. Finally, we added an invitation to others to use our open material (i.e., code) and explore the behaviour of the metrics under other conditions (see top of page 21).

      (3) Decision criteria (strict/lenient/no criterion) need a clearer rationale

      The three continuation rules are a strength of the study, but:

      - The lenient criterion (any negative estimate is considered "beneficial") is unrealistic and should be reframed.

      - The strict criterion (p < 0.025) heavily inflates effect sizes (in Figure 1b) and may distort interpretation.

      It would be helpful to provide a table showing, for each criterion, its real-world analogue (e.g., regulatory requirement, exploratory progression, mechanistic plausibility).

      We have followed your suggestion and added a Table (Table 2) with the description of the criterion and a description of its real-world analogue. No criterion represents an important reference scenario used to evaluate metric behaviour independent of progression decisions. The strict criterion is the closest to regulatory-style evidence. It is also highly selective and therefore might induce biases (e.g., inflated effect sizes). We link lenient to an exploratory decision-making where efficacy evidence is considered in addition to other factors (e.g., safety), but not intended to represent a certain regulatory standard.

      (4) Interpretation of simulation results needs more focus

      The Results section is extremely detailed, making it challenging to identify the central take-home messages. The authors should consider adding a concise summary table comparing metrics on key properties:

      - T1E control robustness.

      - Sensitivity to heterogeneity.

      - Dependence on animal sample size.

      - Dependence on k.

      - Bias under asymmetric effects.

      Moving some nested-loop plot descriptions to the Supplement. Right now, descriptions are technically correct but cognitively heavy.

      We agree with your comment and have attempted to implement it in our summary Table 3, at the end of the results section. We also point readers early on to the Table, so that they can skip the more technical and detailed description if they want (see first paragraph section 3.2, page 12). After some trial and error, we agreed that the chosen columns are the most useful for an applied researcher to get a quick overview. Our table now summarises for each metric its main strengths and weaknesses, its behaviour with increasing heterogeneity, its sensitivity to more animal data (i.e., larger k and larger animal sample size), and its behaviour under effect mismatch (i.e., when the true effect in the animal and human study are dissimilar).

      (5) The discussion should provide explicit recommendations.

      The authors provide high-level recommendations, but the recommendations lack specific guidance. When heterogeneity is low, controlled sceptical p-value works well. When effect sizes differ: weighted Edgington is stable. The authors should avoid using replication BF when the animal effect ≠ human effect. Meta-analysis should not be used when human heterogeneity is high, because of inflated T1E.

      We agree that explicit recommendations would be helpful to the applied researcher. As mentioned in the reply to the previous comment, we have added a summary table which lists the strengths and weaknesses of each metric. We also extended the paragraph in the Recommendations section (on page 21 and 22) to give some examples of scenarios in which certain metrics would be recommended over others.

      (6) Recommendations for applied researchers

      The study is missing an explicit definition of "translation success". The manuscript implicitly defines translation success as: "Both animal and human results show a beneficial treatment effect according to metric X". But this is different from biological translation, which concerns underlying mechanisms. The authors briefly mention this conceptual challenge, but this should be elaborated, as it is central to interpretation.

      Thank you for this comment. We agree that “translation success” was not explicitly defined. We have now added a working definition in the Introduction, clarifying that, in this paper, translation success is defined statistically, and depends on the metric. We now explicitly acknowledge that this is a narrower definition than biological translation. We also elaborate on this distinction in the Discussion where we note that the appropriate metric and interpretation of translation success depends on the translation goal and that statistical translation is distinct from biological translation.

      Minor points:

      (1) The abstract could include a direct sentence on the main conclusion. For example, no metric was uniformly optimal; controlled sceptical p-value and weighted Edgington performed most consistently.

      Our abstract already included main conclusions. We added the word “However” to emphasize the sentence “no metric was uniformly optimal” a bit more.

      (2) The figures are informative, but nested loop plots are very dense. Consider providing a guided example in the figure caption explaining how to read them (as partially done in Figure 1a, but repeat for all).

      We agree that the Figures can be very overwhelming at first. We did not want to add specific helping elements as we did in Figure 1 to not make the figures even busier. The goal was to introduce the reader gently to the nested loop plots via Figure 1 before having them look at the remaining figures. We hope that with the added summary Table and the more detailed recommendations, applied researchers less interested in the statistical details will still find the information most relevant for them easily.

      (3) Methods: Section 2.4 could clearly state that effect sizes are in units of mmHg (blood pressure) from the dataset.

      Thank you for pointing this out. This has been added.

      (4) Results: This section is long; consider adding a brief summary paragraph at the end of 3.2.

      We added a summary table, allowing interested readers to skip the long section entirely.

      (5) Limitations: Add a note about publication bias in animal studies (you mention it in the Introduction, but not in Limitations). Add a statement about effect direction consistency (i.e., animal effect negative but human positive), which is not explored in the simulation grid.

      Thank you for pointing out this inconsistency. A note about publication bias in animal studies was added to the Limitations section (that this was not investigated). A note about opposite animal and human effects was added to Section 2.5 (Simulation conditions) under “Animal and human effect sizes”.

      Reviewer #2 (Recommendations for the authors):

      Animal studies are typically highly controlled, using animal models that are either outbred to provide higher genetic variability or inbred with very little genetic variability and with a specific phenotype. Additionally, many rodent models are incomplete models of the overall human phenotype and are typically used to investigate only one aspect of the condition/disease. Some of the rat animal models that the Terstappen et al. (2020) systematic review used as the simulation parameters for the study included outbred (Sprague-Dawley, Wistar) and inbred Spontaneous Hypertensive Rats (SHR), which have different mechanisms in which hypertensive onset can occur, especially if inducing preeclampsia in outbred animals. Is it feasible to reduce heterogeneity in the animal results if only outbred or only SHR are considered instead? I realize this may reduce the sample size even further.

      You raise an important point differentiating biological (rather than statistical) translation. We have added a sentence about differences between rat models and humans to the new paragraph in the Limitations section (bottom page 20 and top page 21) on the distinction between biological and statistical translation. As for reducing heterogeneity in the animal results by focusing on one type of rats, we agree focusing on one type of rats might reduce heterogeneity. We however consider this reduction to be very small (because the results of the study with SHR are actually comparable to the results with Wistar and SD rats). Therefore, rerunning the simulation would not yield results that differ in any meaningful way from those already reported and the substantial computational effort required to do so is not warranted.

      Reviewer #3 (Recommendations for the authors):

      Overall, I found this a very detailed study. However, my recommendation is to provide a more approachable overview of the results to reach a wider audience. Currently, the article is much more technical and statistically focused. I think two additions could help.

      (1) A summary table of each of the metrics and their strengths and weaknesses under the various conditions (e.g., animal and human study characteristics). Currently, this is done via text, but I think a high-level summary via a table could be a compelling way to make the simulations more approachable for a non-technical audience.

      As requested also by reviewer 1, we have added a summary table.

      (2) Contextualize the findings within the decision-making process a little more. The authors have a well-written limitations section that acknowledges this; however, I think the discussion (and maybe the introduction) could be enriched by putting the simulation findings into context. For example, this paper suggests a framework that includes replication as part of the decision-making process for human trials (https://www.cell.com/med/fulltext/S2666-6340(24)00296-4).

      We agree that situating our metrics within existing translational decision-making frameworks adds important context. We have added a paragraph in the Discussion (before the Limitations section on page 21) clarifying that the metrics evaluated here should not be viewed as standalone decision rules for progression from animal studies to human trials. Several frameworks have recently emerged precisely to guide such decisions in a more structured, multidimensional way. We refer to PATH and also to the GALENOS approach [DOI: 10.1186/s12874-026-02891-4]. Within such frameworks, translation success metrics of the kind evaluated here may provide a quantitative assessment of the consistency between animal and human efficacy findings, thereby informing one component of a broader translational evidence assessment. We have also briefly mentioned at the end of the Introduction (page 4) that frameworks for structuring the use of preclinical evidence in translational decisions are being developed, further motivating the need for quantitative tools such as those evaluated here.

      Below are some additional minor comments for the authors to consider:

      (1) In the abstract (4th line), there is an extra 'l' in failure.

      Thank you for the detailed review. We have fixed this.

      (2) I think since the study is completed, the objectives in the introduction should be past tense, not future.

      We have fixed this.

      (3) The limitations section should include the fixed human sample size. N=107 per group is grounded in the literature, but this varies widely based on the effect size of interest. Again, not material to the point of translation under simulated conditions (of which this would have increased the simulations well above the 648 already included), but given the impact this has on insights, this limits this investigation to a degree and should be acknowledged.

      Thank you for your comment. We have added a note about the human sample size to the paragraph about the simulation conditions in the Limitations section. The human sample size was computed via power analysis as per regulations, but we realize this could change depending on the effect size.

      (4) I appreciate how shrinkage was calculated. Though it is worth noting that the Reproducibility Project: Cancer Biology found much higher rates, which are similar to reports from biotech and pharma (e.g., 11% and 20-25% for Amgen and Bayer).

      We already mentioned the high rates of shrinkage in the Replication Project Cancer Biology (see page 10). We have now also emphasised that one could adapt these levels further depending on the situation.

    1. eLife Assessment

      This fundamental study uses simultaneous EEG and fMRI recordings to shed light on the relationship between alpha and gamma oscillations and specific cortical layers. The sophisticated methodology provides compelling evidence for correlations between oscillatory power and the strength and contents of fMRI signals in different cortical layers. This paper will be of interest to neuroscientists studying the role and mechanisms of alpha and gamma oscillations.

    2. Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two a higher frequency within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state of the art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts is in line with recent findings and theories. In sum, this paper makes a very nice contribution to the literature. In particular, it provides a novel characterization of how oscillatory signals orchestrate the coding of visual contents in the visual cortex and provides a starting point for further research examining how this oscillatory coding changes across visual contents and tasks.

    3. Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. The authors defined both feature-specific and feature-unspecific contrasts, further subdivided by EEG power regressors, to examine how orientation information is signalled across cortical layers. Feature specific contrast was established via comparing trials where stimulus orientation (respectively) was left with those where the stimulus orientation was right. Further specifying it depending on EEG power regressors as congruent where stimulus orientation of EEG regressor matches voxel preference or incongruent (stimulus orientation of EEG regressor does not match voxel preference).

      Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence. On the other hand, the lower-frequency findings are compelling, the functional specificity observed within the alpha frequency band is a particularly interesting result and further highlights the importance of distinguishing feature specificity in order to reveal more nuanced characteristics of neuroimaging data.

      Overall, main findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. The paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings.

      The evidence for gamma involvement in the observed effects is selective: no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 5C,F), with significant effects emerging only in positively responding voxels and only for the contrast between congruent and incongruent conditions in the feature-specific BOLD response. The authors address this in the Discussion, noting that the stimulus may have elicited a weaker gamma response overall, and the contrast of EEG congruent vs. incongruent is necessary in order to achieve the largest contrast-to-noise ratio.

      Authors reported negative relationship between the alpha frequency band and the feature specific BOLD signal increases (for congruent condition, Figure 5A,D). Furthermore, testing for the functional specificity between lower vs. upper alpha (Figure 5B,E), the authors included statistical test on the mixed effects model coefficients, and found significant interaction between alpha frequencies in the congruent condition. This interaction was mainly driven by the upper alpha band (which was later confirmed with simple effects analysis) and revealed stronger negative relationship of upper alpha and the BOLD signal for the subtraction of congruent over incongruent conditions. These are exciting results, which further advocate for a more active role of upper alpha band involvement (relative to lower alpha band) in processing visual features.

      Expanding on this, the authors have also conducted an exploratory analysis of the relationship between the behavioural findings and underlying neural activity for non-oddball trials (Figure S12 in Supplementary Figures). This confirmed a positive relationship between task performance and alpha frequency, suggesting that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power.

      This study provides a valuable and exciting contribution to the literature on oscillatory dynamics and laminar fMRI.

      Comments on revised version.

      Thank you for the thorough revision and for addressing the comments so carefully. The new figures are super beautiful and make the results considerably easier to interpret, they are a real improvement to the paper.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two higher frequencies within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state-of-the-art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts are in line with recent findings and theories. In sum, this paper makes a very nice contribution. I still have a few comments that I outline below, regarding the data visualization, some methodological aspects, and a couple of theoretical points.

      The authors put a lot of effort into the figure design. For instance, I really like Figure 1, which conveys a lot of information in a nice way. Figures 3 and 4, however, seem over engineered, and it takes a lot of time to distill the contents from them. The fact that they have a supplementary figure explaining the composition of these figures already indicates that the authors realized this is not particularly intuitive. First of all, the ordering of the conditions is not really intuitive. Second, the indication of significance through saturation does not really work; I have a hard time discerning the more and less saturated colors. And finally, the white dots do not really help either. I don't fully understand why they are placed where they are placed (e.g., in Figure 3). My suggestion would be to get rid of one of the factors (I think the voxel selection threshold could go: the authors could run with one of the stricter ones, and the rest could go into the supplement?) and then turn this into a few line plots. That would be so much easier to digest.

      We thank the reviewer for their insightful comments. Below we will address each point separately and highlight the changes made to the manuscript. In agreement with the reviewer we have recompiled Figures 4 and 5 (previously Figures 3 and 4). The new figures only present results for the 10% voxel selection threshold (with 5% and 25% moved to Supplementary Figures, see Figures S1-S9). Instead of the radially arranged layout, we opted for a more traditional figure layout, which significantly improved readability.

      (2) The division between high- and low-frequency alpha in the feature-specific signal correspondence is very interesting. I am wondering whether there is an opposite effect in the feature-unspecific signal correspondence. Would the high-frequency alpha show less of a feature-unspecific correlation with the BOLD?

      Following the reviewer’s interesting suggestion, we added the low/high frequency alpha analysis to the feature-unspecific analysis. Indeed, we have found a significant interaction between the sign of the signal change for selected voxel (positive vs. negative BOLD) and alpha sub-band (low vs high frequency alpha). An analysis of simple effects did not reveal any significant effects, however we found a trend level difference (p=0.097) between low and high-frequency alpha for the positive voxel sub-selection. This indicates a stronger negative relationship between upper alpha and the positive BOLD signal as compared to lower alpha. We interpret this result as partial evidence for a feature-related contribution of the upper alpha band. “Active” cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. The negative relationship between alpha and negative BOLD could thus be interpreted as an indirect effect, resulting from a reduced alpha decrease in cortical patches responding to non-attended receptive field locations. However, the involvement of attention-related processes remains speculative, since attention was not explicitly manipulated as part of the experiment.

      We have added Figure 4 B.

      We have also added this section to the Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub><0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      And the following section of the Discussion was extended:

      “The significant interaction between the sign of the BOLD signal deflection and upper or lower α bands (see Figure 4 B) further indicates that multiple α-related processes contribute differentially to positive or negative BOLD. "Active" cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. This hypothesis receives additional support from the trend-level difference in α sub-bands for positive BOLD, indicating that lower α is less related to the active, possibly feature-related processes. The absence of this difference for negative BOLD again indicates a broader, more general process. Future experiments manipulating visual features and attention might reveal a differential upper and lower α response to attended visual features and a more general relationship between α (and possibly superficial layer cortical activity) for suppressed (unattended) receptive fields.”

      (3) In the discussion (line 330 onwards), the authors mention that low-frequency alpha is predominantly related to superficial layers, referencing Figure 4A. I have a hard time appreciating this pattern there. Can the authors provide some more information on where to look?

      We thank the reviewer for pointing out the lack of clarity of this section in the Discussion. We have now rephrased the Discussion, focusing more on the laminar difference and keeping the frequency difference to a separate paragraph. Our main argument for possibly multiple alpha-related processes are twofold: a difference in alpha frequency depending on the underlying analysis (low vs high frequency alpha) and a different layer distribution (superficial layers vs. superficial and deep layers, depending on the analysis). The respective section in the Discussion now focuses on the laminar difference only. We find a negative relationship between alpha and the BOLD signal most prominently in superficial layers (feature-unspecific contrast for the BOLD signal with negative t-values; Figure 4). In addition, we find a superficial and deep layer contribution for the feature-specific contrast (congruent - incongruent; Figure 5A). While the here presented experiment was set out to investigate feature-specific processes, the meaning of the feature-unspecific results are of speculative nature. Future experiments should target the laminar difference between feature-specific and unspecific processes with respect to alpha frequency and layer distribution directly. 

      We have modified the respective sections in the Discussion:

      “Furthermore, we observed that the relationship between the feature-specific BOLD signal and α is predominantly linked to frequencies above 11 Hz (see Figure 5A). An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies. Since individual frequency variations have been included as a random slope in the linear mixed-effects model, these effects cannot be explained by a subset of participants driving lower or upper α separately. Specifically our findings on upper α indicate that α is not exclusively linked to global signal modulations, which has been the traditional perspective [...]”

      “Not only did we find a dissociation in the frequency domain between the relationship of α and the BOLD signal, but furthermore found that the laminar activation patterns provide further evidence for potentially multiple α-related processes. The association between α and the BOLD signal was strongest in superficial layers for negative BOLD activity and feature-specific activity (see Figure 4A and 5A). However, deep layer-related α effects were limited to feature-specific processes only (see Figure 5 A Co-Inco). These findings suggest that superficial layer α reflects are broader, more general process, while deep layer α operates more narrowly, linked to the processing of the visual features themselves. Previous findings using laminar fMRI (which did not include the investigation of oscillatory activity), indicate that superficial layer activity might be more related to the modulation of attention [...]”

      (4) How did the authors deal with the signal-to-noise ratio (SNR) across layers, where the presence of larger drain veins typically increases BOLD (and thereby SNR) in superficial layers? This may explain the pattern of feature-unspecific effects in the alpha (Figure 3). Can the authors perform some type of SNR estimate (e.g., split-half reliability of voxel activations or similar) across layers to check whether SNR plays a role in this general pattern?

      We agree with the reviewer that the vascular draining effect typically leads to increased signal change in superficial layers, the effect on (t)SNR however might be less straightforward. We did not include any counteracting measures, because we were not interested in the amplitude of the signal change, but now include an estimate of tSNR (See Figure S10 in Supplementary Figures). We found that in fact the signal-to-noise ratio is higher in deep layers. Most importantly however, the tSNR layer profiles we identified do not reflect the correlation layer result patterns of the combined EEG-fMRI analysis. This indicates that our results are most likely not the result of tSNR differences. In order to confirm our tSNR pattern we have also conducted a second layer analysis based on the LAYNII toolbox, which assigns voxels between pial and white matter to distinct layers (as compared to our fraction-based approach) and found a similar profile as with our initial analysis. However, absolute tSNR values were found to be higher for our weighted layer analysis. We speculate that while functionally relevant components of the BOLD signal drain towards superficial layers, physiological noise components will drain towards superficial layers as well.

      It is furthermore worth pointing out that for the contrast (congruent - incongruent), the vascular draining effect would cancel out between the conditions. Our findings on superficial and deep layers for those contrasts can hence not be explained by vascular draining at all.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      (5) The GLM used for modelling the fMRI data included lots of regressors, and the scanning was intermittent. How much data was available in the end for sensibly estimating the baseline? This was not really clear to me from the methods (or I might have missed it). This seems relevant here, as the sign of the beta estimates plays a major role in interpreting the results here.

      This is a very important remark and we would like to apologise for the confusion. It was not clear in the manuscript that the GLM was computed on z-transformed fMRI data. We have not specifically collected any “baseline volumes”. A positive beta value would indicate that the sign of the predictor matches the sign of the BOLD signal deflection (and vice versa).

      We have added or modified the following sections in Results and Methods respectively:

      “Before the GLM was computed, the fMRI data was z-transformed across time, separately for each block and voxel.”

      “A general linear model (GLM) has been computed with predictors for each TF bin separately for all voxels in V1 that later have been sub-selected according to the respective condition. Time courses for each voxel have been z-transformed before the GLM was computed for each voxel and experimental block separately. Afterwards, each of the resulting regression coefficients (β values) were multiplied with the voxel-specific layer weights that have been obtained as described above.”

      (6) Some recent research suggests that gamma activity, much in contrast to the prevailing view of the mechanism for feedforward information propagation, relates to the feedback process (e.g., Vinck et al., 2025, TiCS). This view kind of fits with the localization of gamma to the deep layer here?

      (7) Another recent review (Stecher et al., 2025, TiNS) discusses feature-specific codes in visual alpha rhythms quite a bit, and it might be worth discussing how your results align with the results reported there.

      We would like to thank the reviewer for pointing out these papers. Yes, we believe that those could be very related to the effects reported here. At the time of writing the initial manuscript we were not aware of the mentioned publications. 

      We have now included these papers in the Discussion:

      “Recent publications on the information exchange within and between primary visual cortex areas of macaques also reported deep layer γ band activity depending on the stimulus material (Gieselmann et al., 2022; Ferro et al., 2021). Those publications challenge the feed-forward exclusivity of γ altogether by revealing intra-area feedback communication in V1 from layer 5 to layer 6 and layer 6 to supra-granular layers. Possibly, the relationship between γ and deep layer BOLD we observed is also related to similar processes (Vinck et al., 2025).”

      “Similarly, in a recent opinion article, Stecher et al. (2025) promote the idea of "content-aware" α-oscillations. In agreement with our results, the authors argue that α-oscillations are related to content-specific feedback signals, reflected in increased decoding performance based on α power of top-down related processes, even prior to the onset of the stimulus (Hetenyi et al., 2025).. Accordingly, we interpret the lower α effect [...]”

      Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence, but most seem more like a suggestive trend. Nonetheless, the paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings. Further investigation will be important to strengthen and expand upon these insights.

      One of the main strengths of the paper lies in the use of a well-established and straightforward experimental paradigm (the visual oddball task). As a result, the behavioural effects reported were largely expected and reassuring to see replicated. The acquisition technique used is very novel, and while this may introduce challenges for data analysis, the authors appear to have addressed these appropriately.

      Later findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. However, the layer weight calculation is lacking in the manuscript. While it is discussed in the methods, it would help to briefly explain in the results how these weights are calculated, so that the reader can better follow what is being interpreted.

      Line 104 states there is one virtual channel per hemisphere for low and high frequencies. It may be helpful to include the number of channels (n=4) in the results section, as specified in the methods. Also, this raises the question of whether a single virtual channel (i.e., voxel) provides sufficient information for reproducibility.

      We thank the reviewer for encouraging us to clarify the virtual channel selection and we agree that the current description could be misleading. Indeed, we selected 4 virtual channels in total: 1 for each frequency band (alpha/gamma), for each hemisphere separately. The main goal of this selection was to find the clearest response of that frequency band to the task. Previous publications used a supervised (ICA-based) approach to extract those responses. To increase reproducibility, we have chosen an unsupervised beamformer-based approach. The reconstruction of time or frequency-resolved sources in the brain typically yields spatially highly correlated results. Publications focusing on this type of analyses report a spatial extent of typically multiple centimetres, which here is the case as well (see Figure 3A of the updated manuscript). As such, the single voxel selection boils down to selecting the peak response within a large patch of very similarly responding voxels. Using this approach we were able to select the frequency response with the highest possible SNR. We do not however claim that the respective single voxel is exclusively carrying this information. In addition we have added a short explanation to the Discussion, since we believe that virtual channel selection with a different objective (e.g. maximising the difference between conditions or maximising cross-frequency coupling, etc.) could indeed profoundly impact the EEG-fMRI correlation, which would open up opportunities for interesting analyses that are however beyond the scope of this project.

      We have added the following section to the Discussion:

      “Future work might also vary the exact virtual channel selection for obtaining EEG-based regressors. Here, we focused on the grid points (voxel locations) with the strongest α or γ response for each frequency band in each hemisphere, derived from the average frequency response to maximise SNR. However, selecting the respective virtual channels based on the response to specific stimulus features or the interaction between high and low frequency bands are possibilities worth exploring in future work.”

      One area that would benefit from further clarification is the interpretation of gamma oscillations. The evidence for gamma involvement in the observed effects appears somewhat limited. For example, no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 2). Significant effects emerged only when the analysis was restricted to positively responding voxels, and even then, only for the contrast between EEG-coherent and EEG-incoherent conditions in the feature-specific BOLD response. It remains unclear how to interpret this selective emergence of gamma-related effects. Given previous literature linking gamma to feedforward processing, one might expect more robust involvement in broader, feature-unspecific contrasts. The current discussion presents the gamma-related findings with some confidence, and the manuscript would benefit from a more nuanced reflection on why these effects may not have appeared more broadly. The explanation provided in line 230, that restricting the analysis to positively responding voxels may have increased the SNR, is reasonable, but it may not fully account for the absence of gamma effects in V1's feature-unspecific response. Including the actual beta values from Figure 4 in the legend or main text would also help readers better assess the strength and specificity of the reported effects.

      We agree with the reviewer that the missing gamma-band response for the feature-unspecific signal, as well as the limitation of the effect solely to the feature-specific contrast for positive voxel selections only was unexpected. In fact, based on previous literature, we were expecting a feature-unspecific effect in the gamma band as well. However, the literature on laminar level EEG-fMRI is sparse and previous experiments used tasks that did not allow for the separation into distinct features (here left or right-oriented gratings). While we cannot fully explain the absence of the gamma effect for the feature-unspecific condition, we reasoned that our stimuli evoked weaker gamma band responses compared to previous literature. 

      The fact that we only see a significant gamma band response for the contrast for positive voxel selections can be interpreted twofold: First, previous experiments limit their analyses to positive BOLD responses only, for which we find an effect as well. Second, the fact that a significant effect could only be obtained for the contrast, might indicate that gamma band activity is related to the actual features themselves. A cortical column responding to left-oriented gratings would then be related to a gamma band response linked to that orientation. If this response to a single orientation could not be fully captured due to SNR-related issues, we would not see this effect in the congruent-only condition and also not in the feature-unspecific condition (because this boils down to both congruent conditions combined). If gamma-band oscillations are actually reflecting the response of a column to a certain orientation, then the lowest possible response would be found for the exact orthogonal orientation (here the incongruent condition). The contrast between most preferred and most not-preferred orientation might have helped to overcome the inherently low SNR, explaining the results for the contrast.

      Lastly, we did not include actual beta values in the main text, because those might be misleading. We compute the relationship between EEG power and the BOLD signal for every voxel separately, then weighted the result with the respective layer weight and lastly aggregated across voxels.This means that the beta values express the strength of the association between EEG and fMRI for an average voxel. For this reason the values are tiny and the values themselves are less meaningful than “typical” beta values.

      We have added or modified the following sections in the Discussion or Methods respectively:

      “Based on previous literature, we expected a γ band effect for the congruent condition of the feature-specific analysis (Scheeringa et al., 2016), which we did not observe. A possible explanation could be the used stimulus material in our experiment as compared to Scheeringa et al., (2016). Muthukumaraswamy et al., (2013) found that stationary gratings evoke a weaker γ band response as compared to moving annular stimuli that have been used by Scheeringa and colleagues. If γ is related to the processing of the actual features themselves (e.g. to a column preferably responding to left-oriented gratings), then contrasting congruent and incongruent voxel selections provides the largest possible contrast-to-noise ratio (CNR). In turn annular stimuli as previously used might have activated all possible orientations and thus might have greatly boosted γ SNR.”

      “The described procedure of computing a GLM based on z-transformed data using z-transformed predictors yields β-coefficients that reflect the average relationship of a single voxel's BOLD response for a given layer (fraction of the single voxel's β) with EEG power changes of a specified frequency.”

      Relating to behavioural findings for underlying neural activity, could the authors test on a trial-by-trial basis how behavioural performance relates to the BOLD signal / oscillatory activity change? Line 305 states that "Since behavioural performance in the present study was consistently high at 94% on average and participants were instructed to respond quickly to potential oddball stimuli, a higher alpha frequency might reflect a more successful stimulus encoding and hence faster and more accurate behavioural performance." Also, this might help to relate the findings to the lower vs upper alpha functionality difference.

      This is a very interesting suggestion. We now include an exploratory analysis of the relationship between frequency and behavioural performance in the Supplementary Figures (see Figure S12). We did not perform a correlation between behavioural performance and alpha over trials because of the low numbers of oddball trials (N=40) and very limited number of false responses (94% response accuracy on average). However, we computed a correlation across participants. After averaging the alpha time-frequency spectrum across non-oddball trials, the individual alpha frequency was determined by the frequency where the alpha decrease (between 0.1 and 0.8 s post-stimulus) was largest. The correlation between alpha frequency and either reaction times and d’ (as a measure for accuracy), yields a significantly positive relationship between d’ and alpha frequency. This indicates that alpha frequency is related to task performance. We interpret those exploratory findings such that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power. 

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      In Figure 4, the EEG alpha specificity plot shows relatively large error bars, and there is visible overlap between the lower and upper alpha in both congruent and incongruent conditions. While upper alpha shows a positive slope across conditions and lower alpha remains flat, the interaction appears to be driven by the change from congruent to incongruent in upper alpha. It is worth clarifying whether the simple effects (e.g., lower vs upper within each condition) were tested, given the visual similarity at the incongruent condition. Overall, the significant interaction (p < 0.001, FDR-corrected) is consistent with diverging trends, but a breakdown of simple effects would help interpret the result more clearly. Was there a significant difference between lower and upper alpha in congruent or incongruent conditions?

      We thank the reviewer for this important remark and have added a simple effects analysis (see Figures 4 b and 5 b, e). We found that the main driver for the interaction between congruence condition and alpha frequency is upper alpha. Specifically the negative relationship between upper alpha and the BOLD signal is significantly stronger for the congruent condition and weaker for the incongruent condition. This indicates the upper alpha indeed is related to the processing of visual features.

      We have added a simple effects analysis (See Figures 4 and 5).

      We have added or modified the following in Results, Discussion and Methods respectively:

      In Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub> < 0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      “After correcting for multiple comparisons, we found a significant interaction (p<sub>FDR</sub> < 0.001). This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis. We found a significantly stronger negative relationship of upper α and the BOLD signal for congruent selections (p<sub>FDR</sub> < 0.01) and the reverse for the incongruent condition (p<sub>FDR</sub> < 0.01), as well as a significantly stronger negative relationship within the upper α sub-band for congruent over incongruent voxel selections (p<sub>FDR</sub> < 0.01).”

      “This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis, which revealed a significantly stronger negative relationship of upper α and the BOLD signal for congruent over incongruent selections (p<sub>FDR</sub> < 0.001).”

      In Discussion:

      “An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies.”

      In Methods:

      “Significant interactions were decomposed into simple effects using Wald tests on the model coefficients, ensuring that post-hoc comparisons were derived from the same statistical global variance as the primary interaction.”

      Overall, this study provides a valuable contribution to the literature on oscillatory dynamics and laminar fMRI, though some interpretations would benefit from further clarification or qualification.

      Reviewer #3 (Public review):

      Summary:

      Clausner et al. investigate the relationship between cortical oscillations in the alpha and gamma bands and the feature-specific and feature-unspecific BOLD signals across cortical layers. Using a well-designed stimulus and GLM, they show a method by which different BOLD signals can be differentiated and investigated alongside multiple cortical oscillatory frequencies. In addition to the previously reported positive relationship between gamma and BOLD signals in superficial layers, they show a relationship between gamma and feature-specific BOLD in the deeper layers. Alpha-band power is shown to have a negative relationship with the negative BOLD response for both feature-specific and feature-unspecific contrasts. When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency, and can therefore be interpreted as more feature-specific than lower frequency alpha.

      Strengths:

      The use of interleaved EEG-fMRI has provided a rich dataset that can be used to evaluate the relationship of cortical layer BOLD signals with multiple EEG frequencies. The EEG data were of sufficient quality to see the modulation of both alpha-band and gamma-band oscillations in the group mean VE-channel TFS. The good EEG data quality is backed up with a highly technical analysis pipeline that ultimately enables the interpretation of the cortical layer relationship of the BOLD signal with a range of frequencies in the alpha and gamma bands. The stimulus design allowed for the generation of multiple contrasts for the BOLD signal and the alpha/gamma oscillations in the GLM analysis. Feature-specific and unspecific BOLD contrasts are used with congruently or incongruently selected EEG power regressors to delineate between local and global alpha modulations. A transparent approach is used for the selection of voxels contributing to the final layer profiles, for which statistical analysis is comprehensive but uses an alternative statistical test, which I have not seen in previous layer-fMRI literature.

      A significant negative relationship between alpha-band power and the BOLD signal was seen in congruently (EEGco) selected voxels (predominantly in superficial layers) and in feature-contrast (EEGco-inco) selected (superficial and deep layers). When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency than lower frequency alpha. This is interpreted as a frequency dissociation in the alpha-BOLD relationship, with upper frequency alpha being feature-specific and lower frequency alpha corresponding to general modulation. These results are a valuable addition to the current literature and improve our current understanding of the role of cortical alpha oscillations.

      There is not much work in the literature on the relationship between alpha power and the negative BOLD response (NBR), so the data provided here are particularly valuable. The negative relationship between the NBR and alpha power shown here suggests that there is a reduction in alpha power, linked to locally reduced BOLD activity, which is in line with the previously hypothesized inhibitory nature of alpha.

      Weaknesses:

      It is not entirely clear how the draining vein effect seen in GE-BOLD layer-fMRI data has been accounted for in the analysis. For the contrast of congruent-incongruent, it is assumed that the underlying draining effect will be the same for both conditions, and so should be cancelled out. However, for the other contrasts, it is unclear how the final layer profiles aren't confounded by the bias in BOLD signal towards the superficial layers. Many of the profiles in Figure 3 and Figure 4A show an increased negative correlation between alpha power and the BOLD signal towards the superficial layers.

      We thank the reviewer for this important remark. Reviewer 1 raised a similar concern and I would like to refer you to our response to Reviewer 1, point 4. The veinal draining typically results in a higher signal change closer to the cortical surface. We did not take any measures to counteract this effect, but provide an analysis of tSNR in Supplementary Figures (see Figure S10). Possibly due to the drainage of physiological noise towards the surface, we found the highest tSNR in deep, followed by middle and superficial layers. To verify those results we computed the same analysis using a second layering algorithm, which resulted in the same profile, but overall less tSNR. Crucially the tSNR profile is not reflected in our EEG-fMRI results.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      When investigating if high alpha (8-10 Hz) and low alpha (11-13 Hz) are two different sources of alpha, it would be beneficial to show if this effect is only seen at the group level or can be seen in any single subjects. Inter-subject variability in peak alpha power could result in some subjects having a single low alpha peak and some a single high alpha peak rather than two peaks from different sources.

      We agree with the reviewer that a bias in a subset of participants to generally higher or lower alpha frequencies could potentially skew the presented results. While the initially computed model included a random intercept for the frequencies, we have now added the random slope as well. This ensures that the difference between low and high frequency alpha is indeed only driven by the difference in condition and not the result of individual differences across conditions themselves.

      In order to verify that not a small subset of participants is driving the result pattern, we also computed the fraction of participants that either show the dual alpha pattern (i.e. follow the exact pattern of the group average), contribute to the group average with a single peak or contradict the pattern entirely. Thereby, 40.4% of all participants show a dual alpha pattern, 38.4% a single alpha pattern in the direction of the group average and 21.2% contradict the group average. See Author response image 1:

      Author response image 1.

      Alpha Response Patterns with Example Subjects: V1 Feature Specific Contrast

      We would also like to highlight our added exploratory analysis of the relationship between alpha frequency and behavioural performance, which was requested by Reviewer 2, point 3. We find a significant positive correlation between alpha frequency and task performance on a group level. This indicates that higher alpha frequencies might be related to better discrimination of visual features. We speculate that participants with better task performance are capable of modulating their upper alpha more than participants with worse performance.

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      The figure layout used to present the main findings throughout is an innovative way to present so much information, but it is difficult to decipher the main findings described in the text. The readability would be improved if the example (Appendix 0 - Figure 1) in the supplementary material is included as a second panel inside Figure 3, or, if this is not possible, the example (Appendix 0 - Figure 1) should be clearly referred to in the figure caption. 

      Since Reviewer 1 suggested using an entirely different figure layout, we now opted to remove some information from the main text figures (we only show the 10% threshold, but 5% and 25% is in Supplementary Figures) and chose a more common figure layout. See Figures 4 and 5.

      Recommendations for authors:

      Reviewer #2 (Recommendations for the authors):

      The contrasts used in the analysis are not clearly introduced in the main text. While the methods section explains them more thoroughly, some of this explanation would be better placed in the results section, where the contrasts are first used. Specifically, the concepts of "feature-specific" vs. "feature-unspecific" BOLD signals are introduced with a very brief definition, which could be confusing for readers. The same applies to the terms EEG co and EEG inco; it would help to briefly explain these when they are first mentioned in the results. The supplementary figures and legends are helpful, so it is clear that the authors were prioritising clarity overall.

      The respective analyses are now also explained in the Results section:

      “During each trial either a left or a right-oriented grating was presented, from which two types of analyses have been derived: feature-unspecific BOLD activation (i.e. the response to any stimulus orientation), and feature-specific BOLD activation (i.e. the response to a specific stimulus orientation or the contrast between them). Thereby, fMRI data and EEG-based regressors could either be combined congruently (Co) by combining the BOLD signal of orientation-selective voxels with EEG-based regressors built from the same orientation trials, or incongruently (Inco), by combining the orientation-specific BOLD signal with EEG-based regressors built from the other orientation trials. Finally, those two congruency conditions have been contrasted (Co-Inco).”

      Figures are overall clear and illustrative of the results. For Figure 4, however, the use of dotted elements makes it somewhat harder to interpret what's being shown. While the supplementary figure clarifies the findings, rephrasing the figure legend to explain what the dotted lines represent would be helpful.

      Figures 4 and 5 have been replaced with a new layout and legends have been improved.

      The reported ranges overlap (e.g., alpha: 2-32 Hz; gamma: 20-120 Hz). It would be helpful to explain why such overlapping bands were chosen.

      Both frequency bands of interest differ slightly in their later time-frequency analysis (i.e. number of tapers and filter type). The overlap itself is not meaningful per se and results from the selection of a wide band for each respective sub-band. This wide selection was chosen to avoid filter artefacts. For the alpha sub-band, we also wanted to ensure that the beta spectrum is covered which also includes the alpha harmonic and for the gamma band that the full range of high-frequency activity is captured (e.g. EMG activity).

      Only a single time point was used for baseline correction of the low alpha band. Is this typical? The authors note that due to the gradient artefact arising in the pre-stimulus period, the baseline correction is somewhat difficult, although further clarification would be useful here.

      Relatedly, was pilot scanning conducted? If so, was the presence of strong gradient artefacts unexpected? More details about this would strengthen the methodological transparency.

      Indeed only a single time bin was used as the baseline for the alpha sub-band. After the piloting phase a slight adjustment to the final fMRI sequence has been made which was not expected to introduce gradient artefacts so close to the onset of the stimulus. Unexpectedly, those artefacts were visible until 300 ms before the onset of the stimulus. Similarly, a pre-stimulus alpha was observed (starting 250 ms before the onset of the stimulus), which we also aimed to exclude from the baseline period. In the end only the time bin centered at 300 ms prior to stimulus onset was chosen. However, this time bin contains 400 ms of data (the width of the window for the time frequency analysis). Thus, the term time point was misleading, because the actual time window that made up the baseline is 500 ms to 100 ms prior to the onset of the stimulus. 

      We have adjusted our wording in Methods to make this more clear:

      “For this reason, the low frequency baseline period comprised only a single 400 ms time bin centred around -0.3 s, because a pre-stimulus α decrease was expected starting around 0.25 s prior to stimulus onset.”

      Including a one-sentence explanation of the AROS test in the main text for clarity. As line 796 in the methods: "Each significant cluster has been further processed by means of an auto-regressive rank order similarity (aros) test (Clausner and Gentili, 2022). The fundamental idea behind the AROS test is whether group averages (i.e. averages of the signal of cortical layer in the present case), can be ranked such that the rank order is explained significantly better by the data than it would if the average data could not be meaningfully sorted (i.e. is shuffled)."

      An explanation has been added to the Results section:

      “Each significant cluster was then averaged along the frequency dimension at the widest point to enable an auto-regressive rank order similarity (aros) test Clausner & Gentili (2022), testing the laminar activation profile. The aros test transforms the layer averages into a rank order and tests - using a permutation procedure - if the rank order of the layer averages explains the data better than a random rank order (shuffled layer labels) would.”

      Line 223: "In fact, an analysis of the relationship between the EEG signal and the BOLD signal that focused on the feature contrast only (L - R; independent of the comparison to baseline) revealed a trend-level result with an even stronger deep layer contribution as compared to superficial layers." Could you point to which figure represents this finding - Figure 4B?

      This refers to Figure 4A in the old manuscript, for the 25% threshold for the gamma band. Since now the new figures do not include the 25% threshold anymore, it refers to Figure S4i.

      The number of participants is missing from the main text. Including this in the results section would improve clarity.

      The description of our sample has been moved from Methods to Results.

      Given the complexity of the data acquisition and analysis, the well-designed and easy-to-follow analysis pipeline figure (currently in the supplement) would be better placed in the main text.

      The mentioned Figure has been moved to the main text (now Figure 2).

      Also, simply out of curiosity, what do the authors think about the theta blob around 200ms post-stimulus?

      The theta blob most likely reflects the post-stimulus ERP as often observed in response to visual stimuli. We hypothesise that it is stronger in the middle and superficial layers, but we did not want to extend too much the scope of this paper. Additional analyses could be performed in the future on this evoked activity.

      Reviewer #3 (Recommendations for the authors):

      (1) Minor Corrections to the text and figures:

      We would like to thank the reviewer for the very valuable recommendations. Below we shortly describe how each suggestion has been implemented.

      We have made the white box more clear (see Figure 3 B).

      (b) Page 10: Top of 2nd paragraph - 'The full experimental protocol comprised a high resolution anatomical T1 scan lasting for 8 min'. The methods state this scan is 6 min 31 sec.

      The confusion results from the fact that the T1 scan was recorded during a short practice block that the participants performed inside the scanner. This block lasted 8min during which the 6 min 31 sec T1 scan was recorded. We have made this more clear:

      “Once prepared, the participant was placed inside the scanner and performed an 8 min practice block. A T1-weighted scan was acquired during this time in the sagittal orientation using a 3D MPRAGE sequence Brant-Zawadzki et al., (1992) with the following parameters: TR/TI = 2.2/1.1 s, 11° flip angle, FOV 256 x 256 x 180 mm and an 0.8 mm isotropic resolution. Parallel imaging (iPAT = 2) was used to accelerate the acquisition, resulting in an acquisition time of 6 min and 31s.”

      (c) Page 10: 'Stimulus presentation' paragraph - 'Stimuli were projected onto a screen behind the subject's head using'. The use of 'subject' should be replaced with 'participant' throughout.

      We have corrected the phrasing.

      (d) Page 14: Figures 2A and 2B are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (e) Figure 5 caption: 'Regressors are build for each time-frequency bin separately.' should be 'built'

      We have corrected the mistake.

      (f) Page 16, final paragraph: 'Afterwards, each of the resulting regression coefficients (B coefficients) was multiplied with the voxel specific layer weights that have been obtained as described above.' Should be 'were multiplied'

      We have corrected the mistake.

      (g) Page 17: 'Subsequently, separate analyses were done for two frequency of interest (FOI) ranges centerd around' - typo

      We have corrected the mistake.

      (h) Page 17 - 'Within these frequency ranges inferential statistics based a cluster level' - missing word. Should be 'based on a cluster level'

      We have corrected the mistake.

      (i) Page 14 Figure 2B and 5D are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (2) fMRI data pre-processing:

      Please provide a comment on the EEG-fMRI data quality - e.g. tSNR of EPI data. Perhaps example EPI data could be shown in the supplementary information.

      We included the below Figure S11 in Supplementary Figures showing an example EPI. We have also included an illustration of the result of our layering approach. Furthermore, we included a tSNR analysis (see Response to Reviewer 1, point 4).

      On a practical note - with 14-minute long runs whilst wearing an EEG cap, I would expect participant motion to be a concern. Could you provide some metrics on perhaps the average of the mean and maximum per subject displacement/rotation?

      We ensured that participants receive tactile feedback for their respective head motion from a strip of tape span across their foreheads. This resulted in overall manageable motion during each experimental block. During the main experiment, the average framewise displacement was 0.3 mm, with an average total translation of 1.6 mm and an average total rotation of 1.6 deg within each block. 

      We have added Figure S13 to Supplementary Figures.

      We have added a section to Methods:

      “Subject motion per block was low, with a mean (SD) frame-wise displacement Power et al. (2012) of 0.34 mm (0.24 mm) for the main experiment and 0.23 mm (0.22 mm) for the retinotopy (see also Figure S13 in Supplementary Figures).”

    1. eLife Assessment

      The report by Liu and colleagues provides a valuable analysis of environmental adaptation across diverse lineages of the grass Phragmites australis differing by their level of ploidy. The analysis reports solid evidence that lineages with distinct levels of ploidy occupy different climate niches. The use in tandem of regional survey and common garden experiment represents a convincing approach to suggest a correlation between ploidy and climate adaptation. This manuscript will be of interest to a broad community of ecological genomicists interested in how structural variation in gene dosage potentially affects the pattern of adaptation.

    2. Reviewer #1 (Public review):

      Summary:

      The article is testing the relative advantages of plant lineages with differing ploidy and admixture across environmental gradients. The results show that intraspecific variation in ploidy and admixture between lineages impacts plant traits that may enable persistence and range expansion.

      Strengths:

      Suitable marker panel size and convincing results that include attempts to analyse mixed ploidy level data, which is a challenge.

      Weaknesses:

      (1) Inadequate explanation of allele dosage for ploidy levels, some of which do not match the allele counts expected for genome copy number.

      (2) The setup and sample sizes of the common garden experiments are very unclear. The numbers implied are extremely low to draw robust conclusions.

      (3) Unclear how allele dosage is determined. Given it's so central to many analyses, it would be useful to see how this is done rather than use a citation.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript describes a combination of species distribution mapping experimental data from common garden and physiological experiments to project the future distribution of genetic subgroups with the widespread grass Phragmites australis. Overall, the sample sizes seem appropriate for the questions being asked, and the key results regarding projected change in distribution of the focal lineages are well supported. However, at this point, it is difficult to evaluate the broader impact of the work on the field or the utility of the data for the broader community outside of those studying the focal species, P. austrina.

      Strengths:

      A key strength of the paper is the use of common garden and physiological experiments in conjunction with species distribution modeling. The experiments provide a mechanistic basis for the correlations between interspecific lineage and climatic data, suggesting that the distributional patterns are more likely to result from genetic differences rather than limited dispersal among regions. I would, in fact, emphasize the experimental validation of modeling efforts even more in the introduction.

      Weaknesses:

      I see two weaknesses with the framing of the ms and the presentation of the results. First, no data support the claims that polyploidy has any causal effect. The ploidy levels are, in fact, completely confounded with other genetic differences, so it is not possible to eliminate genetic variation, independent of ploidy, as the causative factor. As the authors note, ploidy was not manipulated in the reported experiments. Thus, the focus on polyploidy in the introduction and elsewhere distracts from the novel and informative experiments that were conducted. Second, the manuscript indicates that intraspecific variation is critical for the evolutionary potential of a species to respond to environmental change, but intraspecific variation is seldom considered in species distribution models. To me, an assessment of evolutionary potential requires estimates of heritable genetic variation and responses to selection. The sample sizes presented here are modest to estimate heritabilities, but the manuscript could be framed with this perspective in mind. However, instead, the manuscript performs species distribution modeling on a small number of sub-specific lineages, essentially treating them as homogeneous "species" - thus the analysis commits the same oversimplification that the manuscript highlights, but does so at a finer evolutionary scale than species. Not acknowledging this simplification (or better, examining phenotypic variation within the genetically defined lineages) hinders what would otherwise be a strength of the manuscript.

      The title suggests that asymmetric introgression and thermal tolerance are the most important findings of the work. However, the introduction contains no explanation of the potential importance of gene flow (other than to say that asymmetric gene flow was suggested by some preliminary analyses), and the discussion offers only a limited explanation of either the potential mechanisms underlying the asymmetric gene flow or its importance for the long-term evolution of the species. Similarly, the novelty of combining experiments and species distribution modeling is scarcely mentioned, and there is no exploration of the connection between tolerance alleles and gene flow. Could introgression of heat tolerance alleles alter the spread of the hybridizing lineages, for example? A greater emphasis on these general population genetic parameters could potentially highlight the broader impact of this work.

    4. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback.

      We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them.

      Reviewer #1 raised two important concerns regarding our methodology. First, the determination of allele dosage is insufficiently explained, which is central to our ploidy assignment and downstream analyses. Second, the setup and sample sizes of the common garden experiments are unclear, raising questions about the robustness of our conclusions. We accept these criticisms and will address them as follows.

      Regarding allele dosage, we will add a detailed step-by-step description of our calling pipeline in the Methods section, including the criteria for peak height ratios and thresholds used to assign copy numbers. We will also clarify a crucial biological detail: the common reed (Phragmites australis) is an allotetraploid in its origin. As a consequence, many molecular markers, including the widely used SSR markers in previous studies, behave as disomic markers (i.e., two homeologous copies inherited in a Mendelian manner). Therefore, observing more than two alleles at a locus is indeed indicative of higher-level ploidy (hexaploidy or octoploidy) in this system. We will explicitly state this to resolve any confusion about why tetraploids in our dataset are treated as having a maximum of two alleles, while hexaploids and octoploids can carry more.

      Regarding the common garden experiment, we will explicitly report the replication number for each lineage-by-treatment combination and clarify the experimental design. We will also discuss the statistical approaches used given the sample sizes, while acknowledging that the consistency between experimental results and distributional patterns lends additional support to our conclusions.

      Reviewer #2 raised three substantive framing issues. First, ploidy is completely confounded with genetic background, yet our manuscript places undue emphasis on polyploidy as a causal factor. Second, our species distribution models treat each lineage as a homogeneous entity, failing to capture within-lineage variation and thus repeating the oversimplification we criticize. Third, we insufficiently explore the evolutionary significance of asymmetric introgression, gene flow, and the novelty of combining SDM with experiments. We fully agree with these points and will revise accordingly.

      To address the confounding issue, we will substantially reframe the manuscript to de-emphasize claims about polyploidy as a causal driver, and instead focus on the adaptive differentiation among distinct genetic lineages that happen to differ in ploidy. The Discussion will explicitly state that dissecting ploidy effects from background genetic effects will require future experimental approaches.

      To address the simplification in SDMs, we will add a clear acknowledgment of this limitation, discussing how it may affect predictive accuracy and suggesting that future studies incorporating population-level genomic data could more directly assess evolutionary potential.

      To address the insufficient exploration of introgression and the novelty of our approach, we will expand the Introduction to better highlight the value of coupling controlled experiments with SDMs at the intraspecific level. In the Discussion, we will elaborate on the evolutionary significance of asymmetric introgression, including testable hypotheses about how gene flow might mediate the spread of heat-tolerance alleles and influence lineage geographical limits under climate change.<br /> We also thank the reviewer for the suggestion to emphasize the experimental validation of SDM efforts, which we will incorporate into a revised Introduction.

      Looking beyond the present study, we envision three complementary directions that build upon our current findings. Expanding common garden experiments to include admixed individuals would test whether introgressed genomic blocks confer fitness advantages under thermal stress. Leveraging the population genomic framework established here, we will transition to whole-genome resequencing for selection scans and genotype-environment association analyses to pinpoint adaptive loci and reveal whether heat-tolerance alleles are preferentially transferred via asymmetric introgression. We will also integrate transcriptomic profiling with phenotypic measurements to identify candidate genes whose expression correlates with thermal performance and introgressed ancestry, helping to disentangle ploidy effects from genetic background. Together, these directions span expanded phenotyping, whole-genome resequencing, and transcriptome-guided discovery, forming an integrated framework that moves from the correlative patterns reported here toward mechanistic understanding. These perspectives are briefly outlined in our Discussion, and we hope the present study will serve as a foundation for these future investigations, which we plan to pursue in subsequent work.

      We believe these revisions will substantially strengthen the manuscript.

    1. eLife Assessment

      This work presents a software and hardware suite for targeted photostimulation that can be used in vivo. The package is a well-designed and documented hardware/software suite with a comprehensive build guide. This tool will likely promote important neuroscience advances through targeted real-time perturbation of the cerebral cortex. Overall, this manuscript makes a compelling case on how to design and make available power tools for the research community.

    2. Reviewer #1 (Public review):

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular. The system has struck a reasonable balance between allowing modularity and providing an integrated complete package, but even more flexibility would be welcome for potential users looking to cut costs, as would clearer presentation of such flexibility as already exists.

      Comments and suggestions are mostly minor, as follows.

      (1) Command signals:

      How is the relationship between analog voltage commands and laser power determined? Is this assumed (or required) to be linear (as Figure 7F implies)? Usability and modularity would be improved by an option to measure or provide a calibration curve for systems with a nonlinear mapping between command voltage and laser power.

      For the grid calibration step, how is the initial mapping from galvo voltage commands to image position determined? Presumably, some sort of initial guess or calculation based on the hardware specifications is needed for the grid calibration to be feasible. Also, how are the number of grid lines and the distance between them determined?

      Why is the mapping between analog outputs and hardware (galvos, laser, masking light) fixed? This would be trivial to make configurable and allow labs with existing setups to adopt Zapit without rewiring existing hardware.

      (2) Laser and optics:

      In Figure 1, the authors should consider explaining the scanning principle schematically, i.e., depicting how tilting of the scan mirrors translates via the scan lens into beam displacement in the specimen plane. Perhaps Zemax can be used for accurate rendering.

      Since the unexpanded beam greatly under-fills the back aperture of the lens, the z resolution is presumably terrible - which is good! That is, for the purposes of LSPS, this advantageously avoids focus-dependent effects, which might otherwise arise due to (e.g.) skull curvature. The authors should consider pointing this out, as well as providing an estimate of the z resolution.

      What is the working distance?

    3. Reviewer #2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    1. eLife Assessment

      This important study combines peptide engineering, molecular docking, and functional assays to define the molecular basis of ligand recognition and activation of the human Y4 receptor and to identify three novel small-molecule agonists. The evidence supporting the conclusions is convincing, with complementary experimental and computational approaches providing strong support for the proposed receptor-ligand interactions. While concentration-response analyses of the small-molecule agonists and additional structural or mutagenesis studies would further strengthen the work, these are not essential to support the main conclusions. The work will be of interest to researchers studying GPCR pharmacology, structural biology, and ligand discovery.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes an investigation of peptide analogue agonists selective for the human Y4 receptor for pancreatic polypeptide over Y1, Y2 and Y5 receptors. After studies of mutated Y4R in transiently transfected COS-7 cells, binding models were calculated. Then, screening of a virtual library identified three non-peptidergic (albeit somewhat peptide-like) compounds with potential agonist activity that were subsequently confirmed and furthermore were found to have receptor interactions similar to the peptide analogues. This study provides fundamental new information that improves understanding of the Y4R structure and mechanism of activation by the native agonist and the selective peptide analogues. The non-peptide agonists have potential for future pharmacotherapy.

      Strengths:

      All of the experiments seem to be well performed, using state-of-the-art methods. The manuscript is quite comprehensive and has used a broad range of methods. The conclusions are convincingly supported by the experimental results.

      Weaknesses:

      The mutagenesis was almost exclusively based on the replacement of potentially interesting amino acid residues with alanine. Replacement with other residues, based on modelling and docking, could have refined the model further. Neither molecular dynamics nor cryo-EM was used to study the agonists' interactions with the Y4 receptor and these are therefore likely next steps in the characterization of the Y4R mechanism of activation.

    3. Reviewer #2 (Public review):

      Summary:

      Pelczyk et al. investigated the binding site of the neuropeptide Y Y4 receptor with the aim of identifying novel small-molecule agonists. The authors first assessed small cyclic peptides as tool compounds and then identified interactions between peptides and receptor residues, which were confirmed by single-point mutagenesis combined with functional assays for intracellular signalling. It is interesting that a peptide receptor can be activated by the relatively small cyclic peptides used in the study. The authors identified both common and peptide-specific interactions. The identified interactions guided ultra-large library screening, which yielded 53 compounds, 3 of which were confirmed as Y4R-specific agonists in an IP-one accumulation assay.

      Strengths:

      The combination of techniques (docking, mutagenesis and functional assays) strongly supports the identification and evaluation of small molecules as agonists at the neuropeptide Y Y4 receptor. Functional assays highlight residues that are important for the binding of all tested peptides, as well as residues with peptide-specific importance.

      The structure-activity relationship component of the study nicely highlights which components of the peptide are important for binding to the different members of the neuropeptide Y receptor family.

      Weaknesses:

      It would have been great to see concentration-response curves for the three identified small-molecule agonists, as this would have stengthened the case for these agonists.

    1. eLife Assessment

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

    3. Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The main strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses:

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

    5. Author response:

      eLife Assessment:

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We will make major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We will change the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions will greatly improve our work and presentation and we are grateful to the reviewers and our editors.

      We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will include an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS will strengthen our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results will be reported in a new Supplemental Table.

      We will include a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results will be reported in the revised manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will temper our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We will strengthen our control of clinical confounders, and add clearer statistical correction in the revision. We look forward to conducting future studies to test causality and understand mechanism.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We will note in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single center, cross-sectional. We will make further modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. eLife Assessment

      This is an important study that applies a new chromatin profiling technique to the study of cellular responses to low oxygen. The authors provide convincing evidence for distinct kinetic phases of the response and identify many new putative regulators of the response. This work will be of broad interest to those studying low oxygen responses and transcriptional regulation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have satisfactorily addressed the comments raised in the previous round of review with textual revisions.]

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

    3. Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major comments from the first round of review:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

      Weaknesses:

      There are several areas where the clarity of presentation could be improved:

      (1) Given the importance of the methodology, the methods section needs more detail on how the extent of MNase digestion is chosen to achieve optimal results with MOA-seq. This is described to some extent in the description of control library preparation, but not for the experimental samples.

      We thank the reviewer for noting this unintended omission. We have not updated the Methods section to specify as follows:

      "Digestion patterns were assessed via gel electrophoresis, and the light digest levels ideal for MOA-seq (as per Savadel et al., 2021) were selected as the lightest digest levels that give a pattern of a nucleosomal ladder spanning the entire DNA fragment size range from undigested to mononucleosome bands, as indicated in Figure 1 with the asterisk-marked gel lanes."

      (2) The abstract describes this approach as "native cistrome profiling" but this is misleading since formaldehyde fixation is used.

      We believe the formaldehyde fixation captures native chromatin structure, but indeed we are digesting fixed chromatin and have updated the wording to read as “in situ cistrome profiling.”

      (3) Species- and field-specific jargon and abbreviations need to be clarified on first usage. For example, on page 9: "Downsampling analysis was carried out for two sets of published reference peaks; the CTCF cCRE peak midpoints and for the ERG motif under the ERG ReMap ChIP-seq peaks." The different categories of cCREs were not clearly defined, nor will it be clear what the term ReMap refers to for those outside the field. The sentence after this refers to IDR, which also should be defined.

      We thank the reviewer for highlighting the need for clearer definitions of field-specific terminology and abbreviations. In response, we have revised the manuscript to explicitly define all relevant terms at first mention. Specifically, we now describe the ENCODE candidate cis-regulatory element (cCRE) catalogue and define the individual cCRE categories, including promoter-like (PLS), proximal enhancer-like (pELS), distal enhancer-like (dELS), DNase I–H3K4me3 (K4m3), and CTCF-only regions. We also clarify that ReMap is a curated database of human transcriptional regulator binding peaks derived from ChIP-seq, ChIP-exo, and DAP-seq experiments. Additionally, we now define IDR as the Irreproducible Discovery Rate framework upon first use.

      (4) Figure 4C: Are these motifs examined under MOA sites specifically or anywhere in the genes in question?

      Leading up to and including Figure 4C, we have not yet examined any motifs. Instead, Figure 4C compares gene sets, one defined by our diff-MOA, and those from GO libraries, in this case the "target genes" which are defined by TF-specific studies, primarily ChIP-seq but also related immuno-based mapping techniques. Consequently, the analysis shown in Fig. 4C is not a motif enrichment analysis. Instead, we used the ENRICHR gene set enrichment analysis tool with ENCODE and ChEA consensus transcription factor target gene sets. Thus, the analysis was performed at the gene-set level, and transcription factor motifs were not examined within diff-MOA peaks or elsewhere in the associated genes for Fig. 4C. We note that motif enrichment within diff-MOA peaks was subsequently examined separately in Fig. 6. In Fig. 7, we further examined differentially expressed genes associated with diff-MOA peaks containing enriched transcription factor motifs and used clustering analyses to investigate their regulatory relationships. We have clarified these distinctions in the revised manuscript.

      If the question is about the location of MOA footprints relative to gene structure, we did not examine any MOA sites at any specific location, just overlapping the gene +/- 200 bp, as indicated in Fig. 4B.

      (5) Figure 5B shows that up-DEGs with diff-MOA footprints tend to show more losses of footprints. Do the authors interpret this as a loss of repressor binding?

      Not exclusively, but yes, that is one plausible explanation. That is, the activation (defined by increased RNA levels) via de-repression could be happening. But we also expect these dynamic footprints to be but one component. In other words, we interpret the relationship as consistent with that possibility, but not only that possibility. A logical explanation is that loss of footprint occupancy associated with upregulated genes could be based on displacement of repressive DNA-binding factors, thereby contributing to transcriptional activation. Thus, while loss of repressor binding is a plausible explanation for a subset of these events, additional factor-specific experiments would be required to know for sure in each case. We have added text to the Discussion acknowledging this possibility.

      Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major Comments:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Regarding additional biological replicates, we have acknowledged this point in the discussion. Importantly, we did subject the replicates to IDR analysis, which we explain in the methods as "In accordance with ENCODE ChIP-seq guidelines (Landt et al., 2012), we further evaluated data quality by assessing pooled pseudo-replicate consistency and self-consistency for each individual replicate (Supplementary Table S2)." This IDR analysis demonstrated consistent peaks between our bioreplicates, meeting ENCODE guidelines. In addition, downsampling analysis demonstrated that our sequencing depth of coverage (Supp Fig 1) was over 10-fold greater than required. We do appreciate that it will be useful to have more biological replicates from other cell types, tissues, or organisms, and hope this study prompts just such future research.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

      Live-cell imaging could indeed be interesting, but it is beyond our current capacity to add to this study and consider this an exciting future direction, but presence in the nucleus could include both bound and unbound HIF1, so the results may not easily track the DNA-bound HIF1 only.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In Figure 3B, the x-axis is not labeled.

      Thank you for pointing this out. We have revised Figure 3B by adding the previously missing x-axis label.

      Reviewer #2 (Recommendations for the authors):

      In the abstract, it would be good to define what MOA-seq is and what the cistrome is.

      Thank you for this suggestion. We have revised the abstract to define both MOA-seq (MNase-defined cistrome-Occupancy Analysis sequencing) and the cistrome upon first mention to improve accessibility for readers who may be unfamiliar with these terms.

    1. eLife Assessment

      The manuscript concerns a fundamental and controversial question in Trypanosoma brucei biology and the parasite life cycle, whether or not dividing slender bloodstream forms must transition to growth-arrested stumpy forms before differentiating to the procyclic form in the Tsetse midgut. The authors provide further evidence that slender bloodstream forms can infect Tsetse flies, and that although their differentiation is considerably delayed, they do not become classical stumpy forms during the process. The study is solid in design and execution, and addresses several criticisms made of the authors' earlier work, although discrepancies with results from other laboratories remain.

    2. Reviewer #2 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

      We outline below the 3 points that, in our opinion, remain to be clarified.

      (1) Effect of NAG on slender-form infections in tsetse flies

      The conclusion that "NAG has a negligible effect on slender infections in tsetse flies" based on Figure 1, cannot be fully supported in the absence of a positive control. A relevant positive control is well established in the literature, namely that NAG promotes Tsetse infection by stumpy forms. Without such a control, it is not possible to exclude technical issues (for example, an ineffective NAG treatment), which would yield results similar to those presented in Figure 1.

      We agree that an internal stumpy-form positive control would provide an additional technical reference. However, the enhancing effect of NAG on stumpy-form midgut infections is well established and was also demonstrated under the experimental framework of our original study (Schuster et al. 2021, Figure 2A).

      The purpose of the present Research Advance was therefore not to re-establish the known effect of NAG on stumpy infections, but to test whether slender-form infections require NAG supplementation. Under the conditions tested here, slender bloodstream forms established midgut, proventriculus and salivary-gland infections also in the absence of NAG. We have revised the text accordingly to avoid implying a general absence of NAG effects and to make clear that our conclusion is restricted to slender-form infections under the conditions tested (line 128).

      (2) Infection of non-teneral flies

      Because the experiments shown in Figure 1 (teneral flies) and Figure 2 (non-teneral flies) were not conducted in parallel or under identical conditions, it is important that the figure legends clearly state the parasite numbers used in each case. Specifically, infections of teneral flies were performed with 200 parasites/mL (approximately 4 parasites per bloodmeal), whereas non-teneral infections used 1 × 10<sup>6</sup> parasites/mL (approximately 20,000 parasites per bloodmeal?). At present, this information is scattered across the Methods and Supplementary Tables 1 and 2, making it difficult for readers to immediately appreciate that the parasite load differs by roughly 5,000-fold between these conditions.

      As previously shown by the authors (Schuster et al., 2021) and in the Rotureau laboratory (Tsagmo Ngoune et al.), and as generally expected, the initial parasite dose strongly influences infection outcomes in teneral flies. In this context, it would be informative to know whether the authors have attempted infections of non-teneral flies using lower parasite numbers (noting that Tsagmo Ngoune et al. used a maximum of 10,000 parasites) and what the infection rate was.

      Relatedly, the statement in line 370 appears to be an overgeneralization, as fly age was not directly tested under matched experimental conditions:

      Line 370 - "Here, we unambiguously show that, in the absence of immunosuppressive treatment, slender forms can establish infections in tsetse flies, irrespective of the fly's age or sex."

      We thank the reviewer for highlighting the inconsistent presentation of parasite doses between Figure 1 and 2. We agree this is confusing and have revised the figure legends to clearly state both the parasite concentration (cells/mL) and estimated fly uptake per bloodmeal for each experiment (Lines 143 and 206).

      Regarding experiments with non-teneral flies using lower parasite numbers: We have not tested intermediate doses (e.g., 10,000 parasites/bloodmeal as used by Ngoune et al.) in non-teneral flies. Given that teneral flies already show relatively low infection rates even under optimal conditions, we chose the higher parasite dose (20,000 parasites/bloodmeal) for non-teneral flies to ensure sufficient statistical power for meaningful analysis of infection outcomes across different fly compartments.

      We acknowledge the reviewer's concern regarding the statement in line 370 and have revised this sentence (line 375) to more accurately reflect our experimental conditions, avoiding overgeneralization beyond the specific parameters tested.

      This reads now: “Here, we demonstrate that slender forms can establish infections without immunosuppressive treatment under the conditions tested. This infectivity was observed in both teneral and non-teneral, as well as in both male and female flies, indicating that slender forms retain transmission potential across different fly demographics. However, direct age comparisons under identical parasite doses remain to be tested.”

      (3) Transcriptomic analysis

      Supplementary Figure 8 lacks statistical analysis, which limits its interpretability. Two types of comparisons would be particularly helpful:

      (i) a comparison of PAD1/2 expression levels between slender and stumpy forms at 0 h; and

      (ii) for each gene, a comparison of the overall change in expression (from 0 to 72 h) between infections initiated with slender versus stumpy forms.

      In addition, the figure legend should clarify what "expression levels" refer to. TPM? Normalized counts?

      We appreciate this helpful comment and included statistical analysis for the expression of PAD1 and PAD2 (Supplementary Figure 8) between the two forms for the baseline (0 h) as well as during the differentiation to procyclic forms (0 h to 72 h) by using Welch´s t-test.

      While PAD1 did not show a statistically significant difference in this analysis, PAD2 displayed significant differences in expression dynamics over time. This supports the broader transcriptomic observation that slender- and stumpy-initiated differentiation follow distinct transcriptional trajectories before converging at the procyclic stage.

      We also clarified the figure legends showing the mean log2 counts per million (CPM) values.

      Finally, for the benefit of the field, eLife could encourage publishing a collaborative study in which the Engstler and Rotureau laboratories exchange parasite lines and culture protocols (including media with and without methylcellulose) and perform tsetse fly infections in parallel in their respective laboratories. Such an approach could help resolve the remaining discrepancies and provide a valuable reference for the community.

      We appreciate this constructive suggestion. A collaborative inter-laboratory study in which parasite lines, culture conditions and infection protocols are exchanged between the Engstler and Rotureau laboratories would be a valuable way to address the remaining discrepancies in the field. In particular, parallel infections using matched parasite lines and culture conditions, including media with and without methylcellulose, could provide a useful reference dataset for the community.

      At the same time, such a study would require substantial coordination, reciprocal strain exchange, protocol harmonization and new infection series in two laboratories. It therefore goes beyond the scope of the present Research Advance, which was designed to address the specific methodological concerns raised in response to our original publication. We have restricted our conclusions accordingly and view the proposed collaborative benchmark study as an important direction for future work.

    1. eLife Assessment

      This important study reports the development of the first tankyrase degrader and demonstrates its enhanced ability to inhibit β-catenin signaling compared to conventional tankyrase inhibitors. The evidence supporting the conclusions is comprehensive and convincing, based on rigorous biochemical and cellular analyses. The findings will be of broad interest to researchers studying Wnt signaling, protein degradation, and cancer biology.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation-thereby impairing β-catenin degradation-the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Comments on revised version:

      I had a favorable opinion of the manuscript in the first round of review. I don't have additional comments on the revised manuscript. The manuscript looks fine to me.

    3. Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Comments on revised version:

      I thank the authors for responding to the queries raised in the original review, most of which have now been addressed. This further strengthens this well-conducted study and well-presented manuscript. I congratulate the authors for this interesting and insightful work.

      A few minor points remain:

      I appreciate the authors acknowledge that testing the physical properties of the degradasome puncta is necessary to explore whether they indeed represent condensates. The term "condensates" implies liquid-liquid phase separation (rightly or wrongly). However, this question has not yet been resolved in the case of degradasomes. I therefore suggest the term "condensates" to be avoided. A simple morphological description as "puncta" may suffice.

      I thank the authors for including the additional data comparing tankyrase binding by IWR and IWR-POMA. I agree that using the BRET signal of IWR-POMA is informative. Adding the IC50 values directly to the figure panels (S3E, S3G) would help the reader to quickly assess binding. The comparison between these two panels is insightful.

      Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue?

    4. Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

    5. Reviewer #4 (Public review):

      From the Reviewing Editor:

      This important study reports the development of the first PROTACs targeting the ADP-ribosyltransferases tankyrase 1 and 2, with the goal of inhibiting Wnt/β-catenin signaling more completely than is possible with catalytic tankyrase inhibitors. The work addresses a significant limitation of existing tankyrase inhibitors: although catalytic inhibition stabilizes AXIN1/2 and suppresses Wnt signaling, it also stabilizes tankyrase itself, potentially enhancing non-catalytic scaffolding functions and promoting accumulation of degradasome-like puncta.

      The evidence is convincing. The authors use appropriate and well-validated approaches, including chemical biology, cellular assays, and proteomic profiling, to show that PROTAC-mediated degradation of tankyrase avoids tankyrase accumulation while still stabilizing AXIN and inhibiting Wnt/β-catenin signaling. The data support the conclusion that degradation of tankyrase can separate pathway inhibition from the confounding effects of stabilized tankyrase protein and may therefore offer advantages over conventional catalytic inhibitors.

      A strength of the study is the clear mechanistic comparison between tankyrase degradation and catalytic inhibition. The manuscript provides convincing evidence that the PROTAC and catalytic inhibitors act through distinct mechanisms, with the PROTAC targeting both catalytic and scaffolding roles of tankyrase. The study is well conducted and clearly presented, and the authors have addressed most concerns raised during review.

      A remaining limitation is that the therapeutic potential of the compound is not tested in vivo, for example in APC-mutant colorectal cancer models, APCmin mice, or patient-derived xenografts. Such experiments would strengthen claims about practical efficacy, although they are not essential for the main mechanistic conclusions of the manuscript.

      Overall, this is an important and insightful contribution. It advances the tankyrase and Wnt signaling fields by providing a new chemical strategy to suppress tankyrase function more completely than catalytic inhibition alone, and it offers a useful framework for future therapeutic exploration of tankyrase degradation.

    6. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation - thereby impairing β-catenin degradation - the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Weaknesses:

      (1) A key caveat is that the identified tankyrase degrader also targets GSPT1 for degradation. This raises the possibility that GSPT1 degradation may contribute to the observed β-catenin and tumor growth inhibition.

      (2) The authors address this concern reasonably by showing that DLD1 cells resistant to GSPT1 degradation remain sensitive to the tankyrase degraded.

      (3) To further strengthen this point, the authors might consider generating TNKS1/2 double knockout cells (e.g., in DLD1 or SW480 backgrounds) and demonstrating that the degrader loses its growth-inhibitory effect in these models. However, given the technical challenges of creating double knockouts in cancer cell lines, such experiments could be considered optional.

      We thank the Reviewer for the favorable feedback. The major concern is the collateral degradation of GSPT1. As the Reviewer noted, IWR1-POMA was able to suppress colony formation in DLD-1 cells resistant to a GSPT1/2 degrader (DLD-1R, Figure 6B and S9F), suggesting that TNKS but not GSPT degradation is responsible for growth inhibition.

      We also appreciate that the Reviewer brought it to our attention an important early observation of the TNKS scaffolding effects. Cong reported in 2009 that overexpression of TNKS induced AXIN puncta formation in a SAM but not PARP domain-dependent manner (PMID: 19759537, Ref. 12). We have added this reference to the introduction of TNKS scaffolding in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Weaknesses:

      Although the results are promising and mostly compelling, the claim that the PROTACs provide "a deeper suppression of the WNT/β-catenin pathway activity" requires further corroboration, particularly at endogenous tankyrase levels.

      We thank the Reviewer for the encouraging and insightful comments. The major critique concerns whether TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels. IWR1-POMA reduced the level of cytosolic β-catenin more effectively than IWR1 in Wnt3A-stimulated HEK293 cells without protein overexpression (Figure 1D). IWR1POMA also suppressed STF activity more effectively than IWR1 in DLD-1 cells (Figure S8C) and reduced the expression levels of several WNT/β-catenin targets more effectively than IWR1 (Figure 1G and S8D). These results support that TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels.

      There are also some other points that, if considered, would further improve the manuscript, as detailed below.

      (1) Abstract and line 62: Many catalytic tankyrase inhibitors tend to display toxicity, which is likely on-target (e.g., 10.1177/0192623315621192; 10.1158/0008-5472). This constitutes the main limiting factor for these compounds. An incomplete inhibition of Wnt/beta-catenin signalling may contribute to the challenges, but this does not appear to be the dominant problem. A more prominent introduction to this important challenge is probably expected by the field.

      A previous study showed that G007-LK, a selective TNKS inhibitor, exhibited weak efficacy and dose-limiting toxicity at 5‒30 mg/kg BID or 10‒60 mg/kg QD in various mouse xenograft models (PMID: 23539443, Ref. 28). Similarly, G-631, another TNKS inhibitor, also showed dose-limiting toxicity without significant efficacy at 25‒100 mg/kg QD in mice (PMID: 26692561, Ref. 60). However, other studies showed that G007-LK was well-tolerated at 200 mg/kg QD over 3 weeks in mice (PMID: 29316982, Ref. 61), and treating mice with G007-LK at 10 mg/kg QD over 6 months also improved glucose tolerance without notable toxicity (PMID: 26631215, Ref. 62). Importantly, basroparib, a selective TNKS inhibitor, was well tolerated in a recent clinical trial (PMID: 40964966, Ref. 64), and constitutive silencing of both TNKS1 and TNKS2 for 150 days in APC-null mice prevented tumorigenesis without damaging the intestines (PMID: 31337618, Ref. 8). We have included some discussion of the toxicity issue associated with TNKS targeting at the end of the Discussion section.

      (2) The authors do a good job in setting the scene for the need for tankyrase degraders. Their observations relating to the formation of puncta (degradasomes) being tankyrase-dependent are compatible with a previous study by Martino-Echarri et al. 2016 (10.1371/journal.pone.0150484): simultaneous silencing of TNKS and TNKS2 by RNAi abolishes degradasome formation. The paper is cited as reference 17, but only in passing, and deserves more prominence. (It includes an entire paragraph titled "Expression of tankyrases 1 and 2 is required for TNKSi-induced formation of axin puncta").

      Indeed, Henderson’s 2016 paper (PMID: 26930278, previously Ref. 17, now Ref. 18) shed important light on the role of TNKS scaffolding in the DC. However, whereas this study demonstrated that knocking down both TNKS1 and TNKS2 by siRNA prevented G007-LK to induce AXIN puncta, it concluded that “puncta formation requires both the expression and the inactivation of TNKS,” which is inconsistent with our observations that accumulation of either catalytically active or inactive TNKS can promote AXIN puncta formation. The function roles of TNKS scaffolding in the DC also remained unaddressed. We have included additional discussion of Henderson’s findings in the first paragraph the Discussion section.

      (3) Moreover, the scaffolding concept has been discussed comprehensively in other studies: 10.1111/bph.14038 and more recently 10.1042/BCJ20230230. There are also a few studies that focus on targeting the ankyrin repeat clusters of tankyrase to disengage substrates (10.1038/s41598-020-69229-y; 10.1038/s41598-019-55240-5) that illustrate the concept of blocking the scaffolding function. In that sense, the hypotheses are mature, and it is interesting to see some of them supported in this study. The authors could improve how they set their work into the context of these other efforts and proposals.

      Indeed, Guettler demonstrated in 2016 that TNKS scaffolding could promote WNT/β-catenin signaling, which forms the basis of the current work. Meanwhile, whereas there have been efforts to target the SAM or ARC domain to address TNKS scaffolding by Guettler and Lehtiö, our approach of targeting TNKS for degradation is complementary. We have included in the last paragraph of the Discussion section information on efforts to target the ARC or SAM domains as an alternative approach to suppress WNT/β-catenin signaling without promoting TNKS oligomerization (PMID: 31836723 and 32704068, Ref. 66 and 67).

      (4) In several places in the manuscript, the DC is referred to as "biomolecular condensate", at times even as a "classic example", implying that it operates through phase separation. This has not been demonstrated. In fact, super-resolution microscopy indicates that the puncta are not droplet-like (10.7554/eLife.08022), which would argue against the condensate hypothesis.

      Biomolecular condensates are membraneless cellular compartments formed by phase separation of biomolecules, regardless of their physical/material properties (PMID: 28935776 and 28225081, Ref. 22 and 23). Super-resolution microscopy studies by Stenmark (PMID: 26124443, Ref. 17) showed that AXIN, APC, TNKS, and β-catenin interacted with each other to assemble into membraneless complexes, wherein AXIN and APC formed filaments throughout the DC. Peifer has also summarized evidence that supports the condensate nature of the DC (PMID: 30782412, Ref. 9; see also PMID: 26393419). However, we acknowledge that testing the physical properties of reconstituted DC (for example, PMID: 34352208) with TNKS will provide a better understanding of the nature, for example liquid vs. gel, of these condensates.

      (5) It is beautiful to be able to use IWR1 and IWR1-POMA at identical concentrations for direct comparisons. However, this requires the two compounds to bind to tankyrase similarly well and reach the target to a comparable extent. How sure are authors that target engagement is comparable? Has this been evaluated?

      Using a BRET assay, we have confirmed that IWR1-POMA binds to TNKS1 with affinity comparable to that of IWR1. Details of this study is now included in the Results sections, and the data are presented in the Supplementary Information (Fig. S3E–G).

      (6) Figure 1F: It is not immediately apparent how IWR1-POMA shows more complete containment of Wnt/beta-catenin signalling. Most Wnt/beta-catenin targets lie close to the perfect diagonal, so I do not see how the statement "that IWR1-POMA controlled WNT/β-catenin signaling more effectively than IWR1" (in the legend of Figure 1F) is supported. Minimally, an expanded explanation would benefit the reader. Providing the colour-coding legend directly in the figure would help improve clarity. Also, the panel is very small and may benefit from a different presentation in the figure.

      We have updated Fig. 1F to include an inset of Quadrant III for improved clarity and readability. We have also moved Fig. S7C to the main text as Fig. 1G and added an expanded explanation for these figures.

      (7) Figure 2: The conclusion of a "deeper suppression" of signalling relies on overexpression of tankyrase in an otherwise tankyrase-null background. Have the authors attempted to measure reporter activity or endogenous gene expression without tankyrase overexpression, in Wnt3a-stimulated cells (in the context of a normal Wnt/beta-catenin pathway) or CRC cells at the basal level? Non-catalytic activity in a similar assay has previously been observed upon tankyrase overexpression (10.1016/j.molcel.2016.06.019). Whether or not there is a substantial scaffolding effect at endogenous tankyrase levels after tankyrase inhibition remains unconfirmed, and the PROTAC is a valuable tool to address this important question. The findings presented in Figure S7C and D go some way towards answering this question - these data could be presented more prominently, and similar assays could be performed in other cell systems.

      IWR1-POMA suppressed STF activity more effectively than IWR1 in APC-mut DLD-1 and SW480 CRC cells without TNKS overexpression (Fig. S8C). Similarly, IWR1-POMA provided a deeper suppression of STF signals in HeLa cells transfected with AXIN1 and β-catenin while expressing endogenous TNKS (Fig. 4G). These results suggest that inhibitor-induced TNKS scaffolding plays a significant role at endogenous TNKS expression levels. Following the reviewer’s suggestion, Fig. S7C is now Fig. 1G.

      (8) Line 237/238: "TNKS accumulation negatively impacts the catalytic activity of the DC (Figure 5D)" - the data do not show this. Beta-catenin levels are a surrogate readout for DC function (phosphorylation and ubiquitylation). Minimally, this requires rewording, with reference to beta-catenin levels.

      We have rephrased "TNKS accumulation negatively impacts the catalytic activity of the DC" as "TNKS accumulation negatively impacts the exchange of β-catenin in the DC."

      (9) Line 303-304: Beta-catenin is thought to exchange at beta-catenin degradasomes; this is clear from previous FRAP assays and the observation that phospho-beta-catenin accumulates in degradasomes upon proteasome inhibition (10.1158/1541-7786.MCR-15-0125). However, degradasome size hasn't, to my knowledge, been related to activity. Can this be clarified, please?

      We apologize for confusing β-catenin phosphorylation with β-catenin abundance. Here, we refer the catalytic activity of the DC to as the ability of the DC to promote β-catenin degradation rather than the kinetics of β-catenin phosphorylation. It is commonly observed that AXIN stabilization by TNKS inhibitors increases the DC size and reduces the β-catenin levels. As such, the induction of AXIN puncta by TNKS inhibitors is frequently used as an indicator of WNT/β-catenin pathway inhibition. However, we have found that, TNKS inhibition drives TNKS accumulation, which reduces the ability of the DC to promote β-catenin degradation. We agree that the DC only primes β-catenin but does not catalyze its degradation. We have revised our manuscript as follows: "increasing the local concentration of the DC components improves its 'effective activity'[50,51]."

      (10) There are previous hypotheses/proposals that the sensitivity of CRC cells to tankyrase inhibition correlates with APC truncation or PIK3CA status (10.1158/1535-7163.MCT-16-0578; 10.1038/s41416-023-02484-8). Have the authors considered expanding their cell line panel (Figure S7) to sample a wider range of cell lines, including some that are wild-type with regard to APC or Wnt/beta-catenin signalling in general? This would be a valuable addition to the work. Quantitated colony formation data could be moved to the main body of the manuscript.

      We have so far tested the effects of IWR1-POMA on the proliferation of DLD-1, SW480, HT-29, HCT116, and RKO cells (Fig. 6A and 6B). While a heterozygous Ser45 deletion in CTNNB1 confers resistance to IWR1-POMA, we did not observe sensitivity associated with APC or PIK3CA status. The ability of IWR1-POMA to suppress the growth of RKO cells expressing wild-type APC is consistent with a previous report that knockdown of both TNKS1 and TNKS2 stabilized PTEN to suppress cell proliferation and glycolysis in vitro and tumor growth in vivo (PMID: 25547115, Ref. 48) independently of the β-catenin pathway. We have added this new information as well as quantification of the colony growth results (Fig. S8A, S8B, S9A, S9F, and S9G) to the revised manuscript.

      (11) The manuscript only mentions toxicity (i.e., therapeutic window) in the last sentence of the Discussion section. As this is THE main challenge with tankyrase inhibitors (as mentioned above), can the authors expand their discussion of this aspect? Is there an expectation that PROTACs may be less toxic?

      As discussed above, evidence for on-target toxicity of WNT/β-catenin inhibition is mixed. Yet, the absence of dose-limiting toxicity for basroparib at doses up to 360 mg QD in human (PMID: 40964966, Ref. 64) is encouraging. PROTAC works by catalyzing target degradation, which is different from traditional catalytic inhibitors that require continuous target occupancy at a high level. It remains unclear whether the observed on-target toxicity of TNKSi is associated with TNKS accumulation at high doses, akin to the cytotoxicity induced by PARP1-trapping upon catalytic inhibition. We have included a brief discussion of the toxicity issue in the final paragraph of the Discussion section.

      (12) Figures 3, 4, 5A: For fluorescence microscopy experiments, can these be quantified, and can repeat data be included?

      We have included quantification data and replicate information for Fig. 3–5.

      (13) Figure 4, S6: An additional channel illustrating the distribution of cells (e.g., nuclei, cytoskeleton, or membrane) would be helpful for orientation and context for the AXIN1 signal.

      We have included cell outlines or nuclear staining for Fig. 3, 4, S6, and S7.

      (14) How were cytosolic fractions of cells prepared to assess cytosolic beta-catenin levels? This detail is missing from the methods.

      We have updated the Methods section to include additional details on the preparation of the cytosolic fractions of cells.

      Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

      We thank the Reviewer for the helpful suggestions.

      I have some recommendations listed below for consideration by the authors to strengthen their study:

      (1) The title is slightly misleading, as it is already known that the scaffolding function of TNKS is important within the DC. The authors should consider incorporating the PROTAC targeting aspect in the title (e.g., PROTAC-mediated targeting of tankyrase leads to increased inhibition of betacat signaling and CRC growth inhibition).

      We have modified the title accordingly to "Targeting tankyrase scaffolding in the β-catenin destruction complex by PROTAC overcomes the limitation of catalytic inhibitors in cancer."

      (2) The authors should comment in the manuscript on the bell-shaped curve obtained with treatment of cells with the PROTACs (Figure S2C). This likely indicates tittering of the targets within a bifunctional molecule with increasing concentration (and likely reveals the auto-inhibition conferred by the catalytic inhibition alone).

      As suggested by the Reviewer, the bell-shaped dose-response likely originated from the formation of non-productive binary protein-ligand complexes at high PROTAC concentrations. We have added a sentence to clarify this unique behavior of PROTAC molecules.

      (3) The authors comment that using G007-LK as warehead was unsuccessful, but they do not show data. Do the authors know why this was the case?

      The structure-activity relationship of PROTACs is often unpredictable, as both the kinetics and thermodynamics of target and E3 ligase binding play important roles in promoting efficient target degradation. We have include data on G007-LK based PROTACs (Fig. S2D) in the revised manuscript.

      (4) Throughout the manuscript, the authors need to do a better job at quantifying their results (i.e., the western blots and the IF). For example, the degradation of TNKS1/2 in Figure 1D is not overly convincing. Similarly, the IF data in Figure 3 needs to be quantified in some ways. Along the same lines, the effect of IWR1-POMA treatments on the proliferation of cells and organoids should be quantified using viability assays... There is also no indication of how many times these experiments were performed and whether the blots shown are representative experiments. The quantification should include all experiments.

      We have included quantification of the immunofluorescence images, colony formation data, and Western blots in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) For clarity, can the authors use the official gene names, TNKS and TNKS2?

      We favor using TNKS1 and TNKS2 when referring to the protein for clarity and use TNKS for simplicity when referring to both proteins.

      (2) Line 92: The authors refer to TNKS2 "induction" - it remains unclear what is meant by "induction".

      We have changed "without induction" to "under basal conditions".

      (3) Can the authors please display molecular weight markers for Western blots throughout?

      (4) Line 144: The description "significantly more effectively" refers to Figure S5A, which shows a single, non-quantified Western blot. I don't think significance has been tested, and this statement should be reworded, or quantified aggregate data provided.

      We have added a Supplementary Information file showing molecular weight markers and quantification of Western blots.

      (5) Line 226: "plateaued at a much lower level" - can this be expressed more quantitatively in the text?

      We have included more quantitative information on the FRAP results.

      (6) Line 249: Can the authors repeat the cross-reference to Figure S7A here?

      We have repeated the cross-reference to the figures.

      (7) Line 266: The description of the experiment using the GSPT1/2 degrader CC-90009 would benefit from a brief recap of the purpose as not every reader will be familiar with this common PROTAC off-target. This is a very thorough analysis, though, and commendable.

      We have added background information on GSPT1 degradation to the revised manuscript.

      (8) Figure 1A: Can the number of repeats and the type of repeats be indicated, please?

      (9) Figure 2: Does n refer to biological or technical repeats?

      (10) Figure 5B, D: How many separate experiments are the data based on?

      (12) Figure S3D, S9A, D: number and types of repeats and the nature of the displayed data and error bars need to be included, please.

      (13) Figure S6B, S7B: I can see three data points, but it would still be helpful to state the number and type of repeats in the legend.

      (14) Figures S9A, S9D: There is value in showing the cumulative data from several repeats in the main figure (Figure 6, which currently is only qualitative) rather than the supplementary material.

      (15) Where single Western blots are shown, can the authors indicate how many experiments they are representative of?

      We have included the number of biological repeats for all data.

      (11) Figure S2C: For most graphs, the main response of interest occurs at low compound concentrations. The y-axis scale does not always help the reader to appreciate the effects, as the response seems small against the magnitude of the hook effect. Interrupting the y-axis as in the final panel may help, with y-axis scales consistent over all panels in the figure.

      We have updated Fig. S2C to emphasize on the degradation efficacy.

      (16) The authors may want to give further method details for some of their assays to facilitate replication of their experiments in the future. For example, the STF assay description is currently quite minimalistic. I assume the assay is fairly robust, though. Other details include cell media (general media details and specific additives and their concentrations in the 3D spheroid formation assay), etc. A general look at the methods section will likely be beneficial.

      We have updated the Methods section to provide more detailed experimental information.

      Reviewer #3 (Recommendations for the authors):

      (1) In Figure 2A, one of the most important findings of the manuscript is that IWR1-POMA induced promoted deeper suppression of beta-catenin-mediated transcription. This seems to be the case only at 3.2uM. Is it statistically significant? What are the data points on this graph? What are the error bars?

      We have included statistical analysis as Fig. S5G.

      (2) On Figure 2C and 2D, do the authors know why the TNKS20M1054V mutant is much better at promoting signaling than the TNKS1-PD ? Is it expression levels?

      It is indeed interesting that TNKS2-M1054V promoted significantly stronger WNT signaling than TNKS1-PD. The basis for its strong scaffolding effect is unclear.

      (3) In Figure 4C, the authors claim that when cells are treated with IWR1-POMA, AXIN1 is distributed diffusely throughout the cytoplasm. It appears that small punctae are visible.

      Quantitative analysis (Fig. 4F) suggest that the size of AXIN1 puncta upon IWR1-POMA is rather insignificant.

      (4) Label on Figure 1D has a spelling error TNKS1/2.

      Corrected.

    1. eLife Assessment

      This manuscript presents openretina, an open-source platform that integrates retinal datasets, model training, benchmarking, and in silico analysis tools within a unified framework. The resource is valuable because it addresses long-standing challenges in reproducibility, accessibility, and cross-study comparison in computational retina research, while providing a foundation for community-driven model development and evaluation. The supporting evidence is solid, with the authors demonstrating a functional and well-documented platform across multiple datasets and species, although a clearer discussion of model interpretability, current performance limitations, and data quality standards would strengthen the resource.

    2. Reviewer #1 (Public review):

      Summary:

      This "Tools and Resources" submission describes a platform for the modeling of stimulus-response relationships in the retina. It includes a repository for experimental data sets with standardized programmatic access, and a suite of software for constructing stimulus-response models and evaluating them.

      Strengths:

      (1) The paper is well written.

      (2) The platform could serve an integrative function by connecting different research programs and offering a common baseline for evaluating stimulus-response models.

      (3) The finding that there is "substantial explainable variance remains uncaptured by current models" is a useful insight to motivate further work and measure progress.

      Weaknesses:

      (1) The modeling supported by the package focuses on predictive accuracy at the cost of less interpretability.

      (2) The article needs to make a stronger argument that this style of modeling is fruitful, especially when applied to the retina.

      Main comments:

      (1) Abstract machine learning vs mechanistic models. The "Core + Readout" architecture advocated here seems to be divorced from all the neurobiological detail that is already known in the retina. It mostly aims at prediction, not interpretation. Such a black-box modeling framework is useful in brain regions where we know very little about connectivity, or mechanisms, or even about the primary function being performed there, like in the mammalian cortex. In those cases, any model that can deliver a prediction is a step forward, even if it does not connect to biological mechanisms. But that's decidedly not the situation in the retina, where so many mechanistic details are known: from consensus cell types, to synaptic detail, to single-neuron biophysics, to circuit motifs. How can one connect this ML modeling approach with the extensive mechanistic knowledge available in retinal neuroscience? And can the combination somehow lead to a better understanding? The authors seem to recognize this tension (e.g. line 215ff and 370ff) but don't give it much weight. A stronger case needs to be made here for how this kind of modeling will advance the field.

      (2) The "gradient field" approach. Figure 4c illustrates a case of this dissonance. The gradient field of the response increases with contrast in multiple directions. This is obvious a priori (see line 274) from the more mechanistic model we already have of this On-Off cell. These are the W3 cells described in www.pnas.org/cgi/doi/10.1073/pnas.1211547109. The circuit-based model from that paper, with rectifying on and off subunits from bipolar cells, gives a much more compact explanation for what the neuron does. Because each of the subunits has a spatio-temporal receptive field, this model can predict the entire dynamics to arbitrary stimuli, rather than just 2 dimensions of static stimuli as in the present analysis. So what is the value added here? Again, a stronger case needs to be made that these "Core + Readout" modeling activities enhance understanding.

      (3) The "most exciting input" approach (Line 193ff):

      - Presumably, some power constraint must be put on the stimulus? Otherwise, increasing the contrast will make it more exciting. What are these constraints?

      - Presumably, this optimal stimulus is computed from the model based on non-optimal stimuli? What are the assumptions going into that?

      - The most exciting stimulus is not necessarily the most useful characterization. Near its maximal firing rate, the neuron doesn't discriminate stimuli much, because the slope there is zero (line 237). Instead (or in addition), one would like to know along which stimulus axis the neuron is most sensitive. See e.g. discussion in Dayan & Abbott 2000, Figure 3.11.

    3. Reviewer #2 (Public review):

      Summary

      openretina is a Python package for training and applying convolutional neural network-based models of retinal ganglion cell responses. The package integrates dataloading, model training, and evaluation in a unified framework built on PyTorch Lightning and Hydra, and ships with pre-trained model checkpoints and publicly available datasets (whitenoise, natural scenes) spanning multiple species (marmoset, mouse, axolotl, salamander) and recording modalities (multielectrode array recordings or 2-p calcium imaging). Beyond predictive modelling, openretina includes a suite of in silico analysis tools for probing learned representations, including maximally exciting input synthesis, discriminatory stimulus optimisation, and model weight visualisation. The broader openretina initiative aims to establish a community-driven platform for computational retina research, lowering barriers to entry and facilitating cross-dataset model benchmarking. This is a valuable contribution given the longstanding fragmentation of datasets, codebases, and analysis practices across retina laboratories.

      Strengths:

      The tool has several strengths. By providing a framework built on deep learning infrastructure, the package substantially lowers the barrier to entry for researchers without extensive machine learning backgrounds. The inclusion of pre-trained model checkpoints across multiple species and recording modalities will allow users to apply state-of-the-art models. The in silico toolkit - and in particular the MEI synthesis pipeline - has already demonstrated its scientific potential, with prior work using optimised stimuli to discover a previously uncharacterised RGC type confirmed experimentally, illustrating what becomes possible when these tools are made broadly accessible. The current modular Core + Readout architecture is a well-suited architecture for modeling retina responses. The HDF5-based data standard provides a sensible common format for contributing new datasets. Overall, the initiative is well-motivated, the engineering is competent, and the vision of a collaborative, community-driven platform for retina modelling is one that the retina community would benefit from.

      Weaknesses:

      (1) The in silico tools provided are valuable, but users should interpret their outputs in light of the performance of the underlying models. The predictive performances of current models and datasets in the package are far from performance ceilings.

      (2) The authors appropriately note that optimised stimuli reveal what a neuron responds to but not how the computation is implemented. I would encourage readers to keep this distinction in mind when using the weight visualization tools as well - convolutional filters in a shared, unconstrained core do not map onto retinal circuit elements, and should be treated as model descriptors rather than circuit proxies. For example, RGCs of the same type may appear to sample inputs from two different filters, which should have been a single filter. Or a single RGC may be sampling from two filters, which under more constrained conditions could be approximated with a single filter. These are degeneracies in the CNN modeling framework that should be kept in mind when drawing circuit-level interpretations.

      (3) The datasets currently distributed with the package vary in recording quality, and users should be aware that model performance may not only reflect architectural limitations but may also be limited by noise and data artifacts, including spike sorting errors.

      (4) As the platform grows and community-contributed datasets are added, explicit data quality standards will be essential. I encourage the authors to develop dataset standards to ensure that their resource provides access to highly curated datasets, which I believe is an important step in having high-fidelity models whose functional interpretations can be trusted.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript presents openretina, a Python-based platform designed to facilitate collaborative retinal modeling across datasets, laboratories, species, and recording modalities. The package provides standardized model architectures, evaluation metrics, and analysis tools, while also integrating several publicly available retinal datasets. The authors further demonstrate the platform through examples of in silico analyses and model benchmarking.

      Strengths:

      (1) Emphasis on standardization and reproducibility. Retinal modeling has become increasingly dependent on deep learning approaches, yet datasets and evaluation procedures remain fragmented across laboratories. By providing a unified framework, the authors lower barriers to entry and create opportunities for more systematic comparisons of models and datasets.

      (2) The manuscript is clearly written, and the examples effectively illustrate the range of analyses supported by the platform.

      (3) The benchmarking results are useful, particularly because they reveal substantial remaining gaps between current model performance and explainable variance ceilings.

      Weaknesses:

      Not a weakness per se, but rather a limitation, is that the manuscript focuses on software infrastructure rather than new biological or computational insights. While this is appropriate for a resource paper, some of the scientific examples, such as the gradient-field analysis of ON-OFF cells, function more as demonstrations than as rigorous validations of novel hypotheses. It might be useful to add a few sentences discussing potential scientific projects that can be immediately facilitated by the openretina (the current text in the Discussion focuses more on advancements in the technical/social aspects of science that will be supported by openretina).

      Overall, this is a valuable and timely resource that is likely to benefit the retinal and computational neuroscience communities.

    5. Author response:

      We thank the editors and reviewers for their thoughtful assessment of our manuscript, and for recognizing openretina as a valuable and timely resource for the retinal modelling community.

      We are especially glad that the reviewers appreciated the motivation of the project, the focus on standardization and reproducibility, and the potential of the platform to support systematic benchmarking and community-driven model development.

      We also understand the concerns raised. In the revision of the manuscript, we will strengthen the conceptual discussion of how predictive models, including the current “Core + Readout” models, can contribute to retinal neuroscience alongside more mechanistic and circuit-based approaches. This is a central matter for us, and one that some of us have recently addressed in a broader review on current trends in retina modelling (see https://doi.org/10.1016/j.visres.2026.108854). We will draw on this perspective to better articulate when predictive models are useful, where their limitations lie, and how openretina can provide infrastructure for comparing functional, normative and mechanistic models within a shared framework.

      We will also clarify the scope and limitations of the in-silico analysis methods provided within openretina. This will include a more explicit discussion of how MEIs, gradient-field analyses, and model-weight visualisations should be interpreted.

      Furthermore, we will add more information that will help the reader better judge different aspects of dataset quality, including, for example, spike-sorting or calcium-processing information and explainable-variance distributions. We note, however, that there are many subtle details about experimental workflows that are difficult to capture in compact indicators. In addition, we will make it clearer that the manuscript represents a snapshot of a living resource: The website, dataset cards, documentation, and repository will be the primary source of this information, especially as new datasets are contributed.

      Finally, we will of course address the technical clarifications raised by the reviewers, with the aim of making the manuscript more accessible overall.

      We are grateful for the reviewers’ constructive comments and believe that addressing these points will make our presentation of openretina clearer and more useful to the community.

    1. eLife Assessment

      In this manuscript, the authors describe a cell-specific mechanism by which glutamate transporters regulate the fidelity with which T-stellate cells in the mouse ventral cochlear nucleus relay information from auditory nerve inputs. The study is supported by solid electrophysiological data. It provides valuable insights into how the rapid binding of glutamate to transporters shapes auditory information processing at specific synapses.

    2. Reviewer #1 (Public review):

      In this article, the authors investigate how glutamate transporter function regulates excitability and synaptic coding in T-stellate cells in the mouse ventral cochlear nucleus. They test this in acute brain slices using whole-cell electrophysiology and artificially raise the relative local concentration of glutamate via pharmacological inhibition of transporter proteins. The main finding is that when sub-saturating doses of DL-TBOA are applied, cells become much more sensitive to synaptic input, diminishing the normally high fidelity of EPSP-spike coupling in these neurons. Notably, high-frequency stimulation in the presence of DL-TBOA reveals a large and slowly decaying AMPA receptor component that underlies persistent/rebound firing in earlier recordings. These effects are not seen in other ventral cochlear neurons, suggesting that rapid glutamate clearance in T-stellate cells, particularly, is important for auditory intensity coding. Overall, these experiments are well-performed, and the findings are robust, though there are some aspects that could be expanded to make the work more impactful. These include a better understanding of the relative contribution of neuronal vs glial transporters and an ability to separate the relative contributions of tonic glutamate concentrations in the cleft vs changes in membrane potential in action potential output. Additionally, there were some minor issues of clarity in both the figure presentation and the main text language that should be addressed.

      Major Points:

      (1) Given the dramatic effect of saturating DL-TBOA on tonic leak/RMP and that the sub-maximal concentration used in most of the experiments still varied between 25-50 uM, Figure 1 would be strengthened substantially by a dose-response curve. Ideally, 5 or 6 concentrations, plotting the effect on tonic current or RMP increase.

      (2) Examining the contribution of glial (EAAT1/2) vs. neuronal (EAAT3) transporters (Fig 8) is intriguing but comes across as incomplete here, especially given the small number of recordings. Using a different non-selective EAAT inhibitor (TFB-TBOA) to chase the EAAT1/2 blocker combo seems like an odd choice, given that you have already characterized the effects of DL-TBOA well. One could also try a lower concentration (~50-100 nM) of TFB-TBOA since it is somewhat selective itself for glial EAAT1/2. Given the data presented, neuronal transporters (presumably EAAT3) appear to dominate the rapid clearance of glutamate at this synapse, but this point isn't emphasized or explored sufficiently.

      (3) Separating the effects of depolarization vs. glutamate clearance was never explored. What effect does depolarizing the cell ~10 mV in control conditions (i.e., without TBOA) have on AP number/fidelity during synaptic stimulation experiments? The authors state that submaximal DL-TBOA generally causes no more than a 5 mV change in RMP, but tonic depolarization could also influence spike fidelity. This experiment could demonstrate that the increase in excitability during/after stimulation is not due to increased engagement of voltage-gated channels.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and mechanistically interesting question: whether plasma membrane glutamate transporters contribute only to slow clearance of ambient glutamate or whether they can rapidly shape synaptic signaling during high-frequency auditory activity. This manuscript provides important evidence that EAAT-mediated glutamate uptake is not merely a slow background clearance mechanism but is essential for maintaining reliable synaptic transmission and linear stimulus-intensity coding in ventral cochlear nucleus T-stellate cells during sustained auditory nerve activity.

      Strengths:

      The finding that EAATs may be required for rapid, local control of glutamate during high-frequency auditory nerve activity is interesting and could have broad relevance to auditory processing. The electrophysiological evidence is generally strong, particularly the use of patch-clamp recordings, stimulus trains, partial versus complete EAAT blockade, and comparison with bushy cell/endbulb synapses. The comparison between T-stellate cells and bushy cells/endbulb synapses strengthens the manuscript. The authors demonstrate that EAAT blockade disrupts coding in T-stellate cells but has little effect on bushy cell spike transmission, supporting a cell-type- and synapse-specific role of glutamate uptake.

      Weaknesses:

      However, some mechanistic conclusions, especially the specific contribution of neuronal versus glial EAATs and the absence of glutamate crosstalk between auditory nerve inputs, rely mainly on pharmacological and indirect electrophysiological inference and would be strengthened by additional anatomical, genetic, or direct glutamate-sensing evidence.

      (1) Clarification of DL-TBOA concentration.

      The authors used bath application of 200 µM TBOA and 25-50 µM in the other experiments, stating that "sub-maximal concentrations (25-50 µM)". The authors should provide a clearer rationale for why different concentrations were used across experiments rather than a fixed concentration.

      The reversibility of DL-TBOA effects should be demonstrated by washout experiments. In addition, potential off-target effects of DL-TBOA on postsynaptic receptors, intrinsic membrane excitability, or presynaptic release (e.g., PPR measurement) should be carefully considered. It would also be useful to test the effects of the submaximal DL-TBOA concentrations (25-50 µM) on membrane potential and inward currents, shown in Figure 1, to determine whether these concentrations depolarize the membrane potential in current-clamp mode or induce inward currents under voltage-clamp conditions.

      (2) Potential contribution of altered intrinsic excitability.

      In Figures 3B and 3C, DL-TBOA appears to induce additional action potentials even immediately after the first stimulation, whereas Figures 6 and 7 suggest that the first EPSC is not substantially altered. This raises the possibility that the enhanced firing may partly result from a modest depolarization caused by background glutamate accumulation or from other changes in intrinsic membrane properties after drug treatment. To address this, the authors should provide a quantitative analysis of physiological parameters under submaximal DL-TBOA conditions, including spontaneous action potential frequency, resting membrane potential, input resistance, and spike threshold.

      (3) Spillover/ crosstalk between AN-fiber-synpases.

      The authors should provide more explanation of how altering the number of active auditory nerve fibers demonstrates the absence of glutamate spillover/crosstalk between bouton synapses. Strong stimulation likely recruits more AN fibers, but it may also change release probability, axonal synchrony, or stimulation spread. The authors should more clearly justify the interpretation that strong stimulation recruits additional independent AN fibers rather than altering release probability or activating fibers with different intrinsic properties.

      (4) Interpretation of glial versus neuronal EAAT contributions.

      The authors claim that both neuronal and glial transporters contribute to rapid uptake using pharmacological approaches. The pharmacological data demonstrate that glial EAATs play a major role in glutamate clearance at T-stellate cell synapses. The strong increase in EPSC decay time and synaptic charge after UCPH-101/DHK application supports the conclusion that glial transporters contribute substantially to limiting glutamate accumulation during sustained auditory nerve activity. However, the conclusion that neuronal EAATs contribute directly should be stated with some caution. The evidence for neuronal EAAT involvement is indirect and depends on the pharmacological specificity and completeness of glial EAAT blockade. The conclusion would be strengthened by additional evidence, such as EAAT subtype expression/localization in T-stellate cells or auditory nerve terminals, transporter current recordings, immunohistochemistry, or genetic manipulation of neuronal EAATs. In addition, fitting the decay phase with a double-exponential model may help determine whether glial and neuronal EAATs contribute over distinct temporal windows.

    1. eLife Assessment

      This manuscript describes an important development of several variants of optogenetic tools to control endogenous p53 activity. They are based on peptides competing with Mdm2/MdmX for binding to p53, thus releasing p53 from its negative regulators and stabilizing its cellular levels. In principle, the data are convincing but should be complemented by investigations of p53 target genes at endogenous levels (instead of only reporter constructs). The study therefore remains incomplete but will be of interest to scientists working on optogenetics as well as the p53 field.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors apply the AsLOV2 domain to control the localisation and the exposure of two peptides (PMI and PMI-M3) that compete with Mdm2/MdmX for binding to p53, thus freeing p53 from these negative regulators and allowing its levels to rise. The authors follow an established strategy in optogenetics, which is to combine two layers of regulation for tighter control: (1) caging the peptide into the Ja helix of AsLOV2; 2) sequestration of the peptide away from its site of action using the LOVTRAP system.

      Strengths:

      The authors show that a reporter is activated when cells are exposed to light. A strength is in the lower background that was achieved after adding the second layer of regulation.

      Weaknesses:

      This study claims to be focused on the control of endogenous p53; however, endogenous p53 levels are not quantified. Moreover, endogenous p53 target genes are also not analysed. Only a synthetic reporter is quantified, which has been placed in the genome of HCT116 cells after the creation of a stable cell line. Microscopy images show only one or a maximum of two cells. Finally, the authors claim their strategy is a general one that can be applied to control other peptides, but they do not show this generality in this paper.

    3. Reviewer #2 (Public review):

      The authors developed Opto-MDMi, an optogenetic system for light-controlled activation of endogenous p53. The main idea is to target the p53-MDM2/MDMX regulatory interaction using PMI inhibitory peptides. This is a nice strategy because it avoids overexpression of p53, which can have adverse effects that might confound the study of p53 activity. The authors first tested a LOVTRAP-based localization strategy, which showed some efficacy but also showed basal activation. They then developed a LOV2-PMI peptide-caging module to control the activity of the PMI peptide itself, testing for interactions first in vitro and then in vivo. Finally, they combined the two systems into a dual-lock design, where LOVTRAP controls localization and LOV2-PMI controls peptide activity. This combination led to somewhat more potent stimulation of p53 activity.

      Another useful aspect of the paper is the detailed description of the development and testing of the LOV2-PMI peptide-caging module, which may aid in the design of other LOV2-based peptide-caging designs.

      Strengths

      Overall, the paper is novel and rigorous, and the claims are supported by the data. The optoMDMi tool seems ready for implementation, for example, to manipulate and study the role of p53 signaling dynamics. A few points of clarification would strengthen the work.

      Weaknesses

      The authors develop many tool variants, but there is some lack of clarity over how all of these tools compare to each other, and which ones interested users should use. The work would also be strengthened by showing modulation of endogenous p53 in more than one cell line.

    1. eLife Assessment

      Verma and colleagues interrogate the mechanisms of phagosome maturation arrest during Mycobacterium tuberculosis infection. While cellular events that culminate in this arrest have been largely elucidated, involvement of other organelles, such as mitochondria, has not been highlighted mechanistically. In this valuable study, elements of mitochondrial quality control, such as mitophagy and mitochondrial-derived vesicles involvement, are shown to be paramount in the host-pathogen tussle. The evidence supporting the main conclusions is solid, based on multiple complementary approaches and appropriate controls, although some central mechanistic aspects of the proposed pathway remain only partially resolved.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important and interesting manuscript that uncovers the cross-talk between mitochondrial quality control and phagosome maturation arrest imposed by Mtb.

      A broader host pathogen (intracellular) question pertains to evading phagosomal maturation/arrest. While cellular events that culminate in this arrest have been largely elucidated, involvement of other organelles, such as mitochondria, has not been highlighted mechanistically. This manuscript paints a larger picture than the well-known conventional endolysosomal pathway and portrays a larger landscape involving elements of the mitochondrial quality control, such as mitophagy and mitochondrial-derived vesicles' involvement in the host-pathogen tussle.

      Strengths:

      The systematic characterisation to unravel the interplay between mitochondrial-related pathways and the endolysosomal system allows the authors to unearth some important findings.

      Weaknesses:

      The conclusions drawn require more robust experimentation and analysis.

    3. Reviewer #2 (Public review):

      This manuscript examines the role of autophagy receptor proteins, particularly p62/SQSTM1, in regulating intracellular Mtb survival in human macrophages. Counterintuitively, depleting p62 reduces bacterial survival rather than enhancing it, pointing to a previously unrecognised mechanism. The authors demonstrate that in the absence of p62, mitochondrial quality is maintained through enhanced TOM20⁺ mitochondria-derived vesicle (MDV) biogenesis, dependent on MIRO1/MIRO2. During Mtb infection, these MDVs are redirected to bacterial phagosomes, promoting RAB7 recruitment, overcoming phagosome maturation arrest and facilitating lysosomal targeting of Mtb. In parallel, bacteria experience increased oxidative stress, further contributing to bacterial killing.

      Strengths:

      The mechanistic chain is built using multiple complementary approaches, including genetic perturbation, redox biosensors, metabolic assays and microscopy. The use of primary human macrophages from multiple donors alongside established cell lines increases confidence that the phenotype is not cell-line specific. The replication clock experiment is particularly elegant and clearly demonstrates that the reduction in bacterial burden reflects enhanced killing rather than impaired bacterial replication. Overall, the study identifies an unexpected connection between mitochondrial quality control and phagosome maturation and provides a potentially important advance in our understanding of host-pathogen interactions.

      Weaknesses:

      The study remains entirely in vitro, and the phenotype is absent in mouse macrophages, limiting the immediate physiological and translational relevance of the findings. In addition, many of the central mechanistic conclusions rely heavily on colocalisation analyses, making it difficult to distinguish direct mechanistic relationships from associated trafficking events.

      Overall, this is an interesting and technically strong study that uncovers a novel link between mitochondrial quality control and anti-mycobacterial defence. The mechanistic model is plausible and supported by substantial experimental work. However, several aspects of the proposed pathway require stronger experimental support before some of the broader conclusions can be fully justified.

      Major points

      (1) The central conclusion that TOM20⁺ MDVs are recruited to Mtb-containing phagosomes is based largely on microscopy and colocalisation analyses. Additional orthogonal approaches would strengthen this key aspect of the study and help establish the nature of the vesicles recruited to bacterial phagosomes.

      (2) The proposed mechanism whereby TOM20⁺ MDVs facilitate RAB7 recruitment and reverse phagosome maturation arrest remains incompletely demonstrated. While the MIRO1/2 and RAB7 knockdown experiments support the model, they do not directly establish a causal link between MDV recruitment and phagosomal RAB7 acquisition. Additional experiments addressing this step would considerably strengthen the manuscript.

      (3) The absence of a phenotype in mouse macrophages raises important questions regarding the conservation and physiological relevance of the proposed mechanism. The authors should discuss possible explanations for this species-specific effect and, if feasible, provide additional experimental insight into the basis of this difference.

      (4) The conclusion that mitochondrial quality is maintained despite impaired p62-dependent mitochondrial turnover is based primarily on mitochondrial content, membrane potential, ROS measurements and Seahorse analysis. These are informative but relatively indirect measurements. Additional assessment of mitochondrial turnover by mitophagy would strengthen this aspect of the study.

      (5) The proteins studied throughout the manuscript (p62/SQSTM1, NDP52, OPTN, TAX1BP1 and NBR1) are generally classified as selective autophagy receptors rather than adaptors. The terminology should be corrected throughout the manuscript.

      Minor points:

      (1) Several conclusions throughout the manuscript are based primarily on colocalisation analyses. The limitations of these approaches should be acknowledged explicitly.

      (2) The discussion would benefit from a clearer consideration of how the proposed mechanism relates to established pathways regulating phagosome maturation arrest during Mtb infection.

      (3) The authors may wish to comment on whether enhanced MDV biogenesis could represent a broader host defence mechanism against intracellular pathogens beyond Mtb.

    1. eLife Assessment

      This valuable study provides a cross-species single-cell transcriptomic resource for early female gonadal development in mammals. The data supporting the main conclusion remain incomplete, and experimental validation is needed to strengthen the conclusions. The work will be of interest to reproductive biologists and developmental biologists.

    2. Reviewer #1 (Public review):

      Summary:

      Fang et al. characterize the cellular basis of early ovarian development through a comparative analysis of single-cell transcriptomic data. The authors integrate a novel bovine scRNA-seq dataset, spanning six gestational stages (E38-E112), with stage-matched human (PCW6-16) and mouse (E11.5-E18.5) counterparts. Beyond identifying shared gonadal cell types across these three species, the study uncovers a previously uncharacterized bovine-specific cell population with steroidogenic features. Their analysis highlights conserved, dynamically expressed regulators, including TFAP2C and ZCWPW1 in germ cells and FOS and JUNB in granulosa cells. Furthermore, by employing a machine learning Support Vector Machine (SVM) model, the authors quantify cell-type conservation, demonstrating that while immune and germ cells are highly conserved across species, granulosa cells exhibit substantial evolutionary divergence. This study makes a significant contribution to developmental biology by establishing a comprehensive, cross-species single-cell roadmap of fetal ovarian development. By integrating livestock data with human and rodent models, the authors identify novel cellular states and provide a framework for assessing transcriptional conservation across species.

      Strengths:

      (1) While human and mouse fetal ovaries have been mapped, the inclusion of a high-resolution bovine dataset (107,930 cells total across the study) provides a critical "large mammal" perspective that is often missing from comparative studies.

      (2) The identification of a bovine-specific cell population is an important finding. It suggests that ruminants may have a different developmental timeline for steroidogenic precursors (potentially theca cell ancestors) compared to rodents or humans.

      (3) Training a Support Vector Machine (SVM) to quantitatively assess cell-type similarity is a major strength. It moves beyond qualitative UMAP "eye-balling" to provide a statistical probability of conservation.

      (4) The study links gene expression to higher-order biological processes like epigenetic reprogramming and cell-cell communication (CellChat), providing a holistic view of the gonadal niche.

      Weaknesses:

      (1) The authors integrated publicly available scRNA-seq datasets generated across different laboratories and technical platforms. However, the specific methods used to control for and evaluate batch effects are not clearly described. It is critical to clarify whether the observed species-specific differences are purely biological or partly influenced by technical variation between datasets.

      (2) A challenge inherent to all single-cell studies is the reliance on manual marker-gene-based annotation. While this is standard practice, it remains unclear how robust these assignments are, particularly for the novel "bovine-specific" population. Further evidence or cross-validation (e.g., through varied clustering resolutions or automated annotation tools) is required to ensure these clusters represent true biological states rather than computational artifacts.

      (3) The authors utilized a linear SVM to assess cross-species similarity. However, it is not clear how this model performs compared to established single-cell mapping and comparative tools (e.g., MetaNeighbor or Seurat v5). Providing a justification for this specific SVM-based approach, or a brief comparison with existing benchmarks, would strengthen the methodological rigor of the study.

      (4) While the computational evidence is compelling, the study would be significantly enhanced by independent validation of the "unclassified bovine-specific" cell population. To confirm the biological reality and reproducibility of this novel cell state, the authors should provide additional evidence. This could include in situ validation (e.g., immunofluorescence or in situ hybridization) to determine its physical location and morphology within the gonad, or demonstrating the presence of this specific cell population within an independent, non-overlapping bovine dataset.

    3. Reviewer #2 (Public review):

      Summary:

      The authors generate a comparative single-cell transcriptomic atlas of fetal ovarian development in cattle, human, and mouse, with the goal of identifying conserved and species-specific cellular and molecular features of early ovarian differentiation. The study provides a valuable resource for the field and reveals potentially interesting species-specific characteristics, including a putative bovine steroidogenic cell population. While the dataset is substantial and the computational analyses are generally appropriate, several major conclusions rely primarily on computational inference without independent experimental validation, limiting the strength of evidence supporting some of the central claims.

      Strengths:

      This study provides a valuable cross-species single-cell atlas of fetal ovarian development by integrating newly generated bovine data with human and mouse datasets. The work fills an important gap in reproductive biology and offers a useful resource for investigating conserved and species-specific features of ovarian development.

      The analyses are comprehensive and combine developmental trajectory reconstruction, regulatory network inference, cell-cell communication analysis, and cross-species classification. The identification of a putative bovine-specific steroidogenic cell population is particularly intriguing and may provide a basis for future studies of species-specific ovarian development.

      Weaknesses:

      The main limitation is that several key conclusions rely primarily on computational analyses without independent experimental validation. In particular, the proposed bovine-specific steroidogenic cell population, which represents the major novel finding of the study, is supported only by transcriptomic evidence.

      In addition, many mechanistic interpretations derived from trajectory, regulatory network, and cell-cell communication analyses remain speculative. While the study succeeds as a comparative resource, the evidence supporting several of the central biological claims remains incomplete, and the biological significance of some cross-species differences is not fully explored.

    1. eLife Assessment

      The Review Article by Bal and co-workers presents an overview of skeletal muscle physiology, focusing on Sarcoplasmic Reticulum and Mitochondria-Associated Membranes (MAMs) in relation to calcium handling, ROS, and signaling. It provides a foundation based on the current literature but could have gone further by identifying future research directions and potential avenues for therapeutic intervention.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript is a narrative review addressing age-related alterations in SR-mitochondria interactions in skeletal muscle and their contribution to sarcopenia. It synthesizes existing literature on calcium signaling, mitochondrial dynamics, redox balance, and structural remodeling, and discusses potential interventions including exercise and pharmacological strategies. While the topic is timely and relevant, the manuscript largely reiterates established concepts without providing sufficient conceptual novelty, critical synthesis, or mechanistic insight beyond the current literature.

      Strengths:

      (1) Timely topic: The focus on SR-mitochondria communication in aging muscle is relevant and of growing interest.

      (2) Broad coverage: The review compiles a wide range of literature spanning calcium handling, mitochondrial biology, ROS signaling, and exercise physiology.

      (3) Clear organization: The manuscript is structured logically with thematic sections (SR, mitochondria, MAMs, aging, interventions).

      (4) Didactic value: Could serve as a general overview for non-specialists entering the field.

      Weaknesses:

      (1) Lack of novelty and conceptual advance: The manuscript does not offer new hypotheses, frameworks, or critical reinterpretation of the field. Most statements summarize already well-established knowledge, and no unifying model or novel perspective is developed to justify publication in a high-impact journal like eLife.

      (2) Limited critical analysis: The review is predominantly descriptive rather than analytical. Conflicting findings (e.g., MFN2 roles, MAM density changes, Ca²⁺ overload vs deficiency) are mentioned but not critically evaluated or reconciled. There is little discussion of limitations in the cited studies or gaps in the field.

      (3) Overgeneralization and speculative claims: Several assertions are presented with insufficient nuance (e.g., causal links between MAM disruption and sarcopenia, or therapeutic efficacy of interventions). The distinction between correlation and causation is often unclear, reducing scientific rigor.

      (4) Insufficient depth for a specialist audience: Despite its length, the manuscript lacks mechanistic depth in key areas (e.g., precise molecular regulation of MAMs in vivo, tissue-specific differences, quantitative aspects of Ca²⁺ flux). It reads more like a textbook summary than a high-level scholarly review.

      (5) Redundancy and verbosity: Many sections repeat similar concepts (Ca²⁺ dysregulation, ROS effects, mitochondrial dysfunction) without adding new insight, leading to an unnecessarily long manuscript with limited added value.

      (6) Weak integration of recent literature into a coherent narrative: Although many references are cited, they are not effectively synthesized into a cohesive argument. The manuscript lacks a strong central thesis or clearly defined take-home messages.

      (7) Limited translational or experimental perspective: The section on therapeutic targeting is largely speculative and does not critically assess feasibility, limitations, or current clinical evidence.

    3. Reviewer #2 (Public review):

      This review addresses a highly relevant and timely topic, namely the role of sarcoplasmic reticulum-mitochondria communication and mitochondria-associated membranes (MAMs) in skeletal muscle aging. The manuscript successfully brings together literature from several interconnected fields, including calcium signaling, mitochondrial biology, excitation-contraction coupling, muscle metabolism, and sarcopenia. Given the growing interest in organelle crosstalk as a determinant of muscle health and disease, the topic is undoubtedly of considerable interest to the readership and has the potential to make a valuable contribution to the field.

      However, in its current form, the manuscript devotes a substantial proportion of its content to the description of well-established concepts that are already extensively covered in the literature. Large sections are dedicated to general skeletal muscle physiology, excitation-contraction coupling, calcium handling, mitochondrial biology, and MAM structure and composition. While this background information is useful, the level of detail is often excessive for a review that aims to focus on aging-induced alterations in SR-mitochondria interactions. As a consequence, the central theme of the manuscript becomes diluted, and the review reads more like a broad overview of skeletal muscle physiology than a focused analysis of aging-related MAM remodeling.

      In contrast, the sections specifically dedicated to aging and MAM dysfunction, which represent the most novel and potentially impactful aspects of the review, are comparatively brief and largely descriptive. The discussion of how aging alters MAM architecture, calcium microdomains, mitochondrial calcium signaling, and organelle communication would benefit from substantially greater depth. For example, although the manuscript highlights alterations in proteins such as MFN2, IP3R, VDAC, and MCU, the mechanistic implications of these changes for sarcopenia and age-associated muscle dysfunction are not critically developed. Similarly, the review would be strengthened by a more comprehensive discussion of the evidence linking MAM disruption to impaired muscle performance, metabolic inflexibility, denervation, and mitochondrial dysfunction during aging.

      Another limitation is that much of the manuscript summarizes published findings without sufficiently evaluating the strength of the evidence or discussing existing controversies. Several statements imply causal relationships between MAM disruption and sarcopenia, whereas in many cases, the available data remain largely correlative. The authors should more clearly distinguish between established mechanisms, experimental observations, and emerging hypotheses. A more critical assessment of conflicting findings, particularly regarding the role of MFN2 and the dual consequences of altered mitochondrial calcium uptake, would considerably improve the scientific rigor of the review.

      A major omission concerns the role of mitochondrial Ca²⁺ uptake in skeletal muscle physiology and aging. Throughout the manuscript, mitochondrial Ca²⁺ uptake is presented as a central determinant of muscle function and as a key mechanism linking MAM disruption to sarcopenia. However, the authors do not adequately discuss evidence that challenges this view. In particular, genetic mouse models lacking MCU exhibit surprisingly mild skeletal muscle phenotypes under basal conditions despite a near-complete abolition of rapid mitochondrial Ca²⁺ uptake. These findings have generated considerable debate regarding the physiological importance of mitochondrial Ca²⁺ uptake for muscle function and metabolic regulation. While MCU deletion clearly affects exercise adaptation and certain stress responses, the relatively modest baseline phenotype suggests the existence of compensatory pathways and raises important questions regarding the extent to which impaired mitochondrial Ca²⁺ uptake alone can explain age-associated muscle dysfunction. A balanced review should acknowledge these observations and discuss the ongoing debate regarding the relative contributions of mitochondrial Ca²⁺ deficiency versus mitochondrial Ca²⁺ overload in aging skeletal muscle.

      Similarly, the discussion of MFN2 would benefit from greater nuance. The manuscript largely presents MFN2 as a structural tether linking the sarcoplasmic reticulum and mitochondria. However, MFN2 is a multifunctional protein with well-established roles in mitochondrial fusion, mitochondrial network organization, mitophagy regulation, and metabolic signaling. Consequently, many of the phenotypes associated with altered MFN2 expression cannot be unequivocally attributed to changes in MAM formation. The review does not sufficiently distinguish between the effects of MFN2 on organelle tethering and its effects on mitochondrial dynamics. This distinction is particularly important because several studies have questioned whether MFN2 acts primarily as a positive tether, a negative regulator of contacts, or whether its influence on organelle communication is secondary to its role in controlling mitochondrial morphology. As a result, attributing age-related alterations in SR-mitochondria communication solely to changes in MFN2-mediated tethering may oversimplify a considerably more complex biological scenario.

      The manuscript's organization could also be improved. The sections discussing aging-related alterations, mitochondrial dysfunction, calcium dysregulation, oxidative stress, and therapeutic interventions contain significant overlap and repetition. Streamlining some background sections and reallocating space to a more detailed discussion of aging-specific mechanisms would help maintain focus and improve readability. In particular, the manuscript would benefit from expanding the sections on aging-induced MAM remodeling, age-dependent changes in MAM composition and ultrastructure, and the potential of MAM-targeted interventions as therapeutic strategies for sarcopenia.

      Finally, the review would gain from a stronger future perspectives section. Several important questions remain unresolved, including whether MAM disruption is a primary driver of muscle aging or a secondary consequence of mitochondrial dysfunction, how MAM architecture differs among muscle fiber types during aging, and whether MAM-associated proteins could serve as reliable biomarkers or therapeutic targets in human sarcopenia. Highlighting these knowledge gaps would further enhance the review's impact.

      Overall, the manuscript covers an important and emerging area of research and contains a valuable compilation of the relevant literature. Nevertheless, substantial revision is required to reduce the emphasis on well-established background information, deepen and critically analyze the aging-specific sections, and provide a more focused discussion of the role of MAMs in skeletal muscle aging and sarcopenia.

    1. eLife Assessment

      This study investigates the cellular mechanisms underlying theta-nested gamma oscillations in the medial entorhinal cortex; the experiments are rigorous, and the analyses and modeling provide potentially useful insights into cell-type-dependent circuit dynamics. However, the evidence supporting several key conclusions remains incomplete. The study is limited by conceptual constraints in the experimental design and a modeling approach that does not fully address underlying physiological mechanisms. Overall, this is a careful study that addresses how distinct neuronal populations in superficial MEC participate in theta-gamma coordination and provides new data linking cell-type-specific activity patterns to oscillatory network structure.

    2. Reviewer #1 (Public review):

      Summary:

      The question posed on cell-type-dependent relationships to theta-nested gamma rhythms is an important one. The authors use a variety of ontogenetic, imaging, electrophysiology, and computational techniques to show that reciprocal interactions between excitatory neurons and interneurons in the medial entorhinal cortex generate gamma oscillations. They measure LFP gamma, gamma power of postsynaptic currents in different neurons, spike phases with reference to LFP gamma, and spatial correlations of membrane potentials across a large population of neurons. Arguing (correctly) that gamma rhythm in this setting is generated through a pyramidal-interneuron network gamma (PING) mechanism, they demonstrate cell-type-specific differences in gamma phase-locking. While they show spatial dependencies of sub-threshold voltages and even argue for topographic clustering, these could simply be reflections of the synchronous stimulation paradigm that they use.

      Overall, I appreciate the methodology and rigor, but would have expected more from the study in terms of relevance to physiological stimulation conditions as well as in terms of mechanisms underlying the differences that they report here..

      Strengths:

      The authors are rigorous in how they conduct the experiments, report the data, and perform the analyses. The modeling respects the heterogeneities and is truthful to the experimental design. The conclusions on PING mechanisms are fine, but are not unexpected given the circuitry of the mEC.

      Weaknesses:

      The interpretation of the conclusions, while for the most part is fine, could have been better, especially given the conceptual limitations of the experimental design. The modeling part could have gone beyond simple descriptive matching and addressed mechanistic questions.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors studied the cellular mechanism of theta-nested gamma oscillations in the medial entorhinal cortex (MEC) in vitro. The theta-nested gamma activity was induced by theta-modulated optogenetic stimulation of CaMKII+ neurons. In Figures 1 through 4, they describe the firing phase, synaptic input, and LFP-IPSC coupling of stellate cells, pyramidal cells, and interneurons. They then conducted voltage imaging, capturing the simultaneous activity of 41 cells, and found that subthreshold membrane potentials cluster in a weakly distance-dependent manner (Figure 5). The experiments and analysis are done rigorously for the most part.

      However, the results described in Figures 1 to 4 are largely descriptive and highly similar to those in their recent publication, which utilized almost identical experiments. While the voltage imaging data during theta-nested gamma oscillations are novel, the authors report data from only a single experiment, leaving it unclear whether the results are reproducible. Furthermore, without a comparison to in vivo data, it remains unclear what novel insights this manuscript provides to advance our understanding of the cellular mechanisms underlying theta-nested gamma oscillations.

      (1) The authors recently published another paper on the topic of theta-nested gamma oscillations in the MEC (Williams et al., eNeuro, 2026). In that study, they utilized a Thy1 promoter instead of the CaMKII promoter used here. The motivation for testing the CaMKII promoter in the current manuscript, as well as the novel insights expected from this experimental setup, remains unclear. Given that existing literature suggests inhibitory MEC cells play a critical role in theta activity (e.g., Gonzalez-Sulser et al., 2014)-implying that theta modulation should drive inhibitory rather than excitatory cells-the previous use of the Thy1 promoter appears closer to in vivo conditions than the CaMKII promoter used here.

      The overall conclusion of the current manuscript is that excitatory-inhibitory (E-I) interactions dominate the generation of theta-nested gamma oscillations. However, in their previous eNeuro paper, the authors demonstrated that the interneuron network gamma (ING) mechanism can sustain gamma oscillations without excitatory synaptic transmission. It seems expected that excitatory cells would be involved when the optogenetic stimulation selectively drives excitatory cells. If CaMKII stimulation is less physiological and artificially forces the theta-nested gamma activity to rely on excitatory connections, this conclusion could be misleading. It may potentially describe a mechanism that is irrelevant to physiological processes in vivo. Please see my comment 3, which is related to this point.

      In addition, Figures 1 and 2 heavily overlap with the authors' previous eNeuro publication. The differences in experimental settings and the motivation for performing almost identical experiments must be clearly articulated prior to these figures to avoid confusion. The authors must also justify why it is necessary to present such similar data, and explicitly point out the novel findings in the current paper compared to their previous work.

      (2) Using voltage imaging to investigate theta-nested gamma oscillations is novel. However, the impact of the findings from this experiment appears minimal in the manuscript's current state. The most novel and interesting observation is likely presented in Figure 6, where the authors identified clustered voltage correlations. However, this appears to be an n=1 experiment, and these findings should be replicated at least in a few experiments. Furthermore, the manuscript lacks a discussion or interpretation of this observation, making it unclear whether the result is biologically meaningful. Please find specific suggestions regarding this point below.

      (3) The authors' primary motivation for investigating the mechanisms underlying theta-modulated gamma oscillations is their potential role in grid cell firing. Therefore, it is critical that the mechanisms studied here in vitro accurately reflect in vivo processes. For this reason, greater effort should be made to better link this in vitro study with existing in vivo data. Numerous public in vivo datasets are available that detail the firing activity of putative principal cells and interneurons during exploratory behavior in mice. Intracellular recordings in awake animals have also been published, some of which the authors already cite. The data presented in Figures 1 and 4, for example, could be straightforwardly compared with those existing in vivo metrics. Furthermore, available in vivo silicon probe recordings could provide a reliable estimate of the spatial distribution of gamma-related spike activity. Such data should be compared with the voltage imaging results presented in this study.

      This limitation connects back to the first point. In this manuscript, the authors tested a different method for inducing theta-nested gamma oscillations (via the CaMKII promoter) than in their recent eNeuro paper (via the Thy1 promoter). The outcomes of these two induction methods must be systematically compared against in vivo data to determine which approach aligns more closely with physiological conditions. Without such a comparison, the scientific justification for testing a different promoter in this study remains unclear.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Williams et al. combine optogenetics, whole-cell electrophysiology, local field potential recordings, large-scale voltage imaging, and computational modeling to investigate the cellular and circuit mechanisms underlying theta-nested gamma oscillations in superficial medial entorhinal cortex (mEC). The authors propose that fast-spiking interneurons receive strong gamma-frequency excitatory drive and provide rhythmic inhibition onto principal neurons, supporting a pyramidal-interneuron network gamma (PING) mechanism. They further report cell-type-specific differences in gamma phase locking, spatial clustering of subthreshold voltage signals, and a network model reproducing several observed features, including interneuron bursting and gamma-cycle skipping in excitatory neurons.

      Strengths:

      The study is technically sophisticated and addresses an important question in entorhinal circuit function. The combination of intracellular recordings, voltage imaging, and computational modeling is a clear strength.

      Weaknesses:

      Several key conclusions developed from experimental results require additional raw data, statistical support, clearer methodological description, and more cautious interpretation. The computational modeling focuses primarily on stellate cells, whereas the experimental results suggest an important role for pyramidal neurons in PING dynamics. This creates inconsistency between theory and experiments.

    5. Author response:

      We thank the editors and reviewers for their thoughtful comments. Below, we list our provisional responses to the reviewers’ major points:

      On the rationale for CaMKIIα versus Thy1-driven stimulation and physiological relevance: We agree that we did not make clear the motivation for using CaMKIIα-driven stimulation, distinct from the Thy1-driven paradigm in our previous work (Williams et al., 2026). Using the Thy1 driver, both excitatory and inhibitory cells received direct theta drive. In contrast, CaMKIIα expression is largely restricted to principal neurons. Comparing these models lets us isolate a "driven I-cell" PING mechanism from the "E cell recovers first" mechanism relevant when interneurons are also directly driven.

      Regarding physiological relevance, Gonzalez-Sulser et al. (2014) found that septal GABAergic projections selectively and directly inhibit mEC interneurons, rather than exciting either principal cells or interneurons, implying that theta drive in vivo likely acts through rhythmic disinhibition of interneurons rather than direct excitation of any cell type. Neither the Thy1 nor the CaMKIIα paradigm reproduces this disinhibitory mechanism: both rely on excitatory optogenetic drive rather than rhythmic inhibition of interneurons, and replicating the natural drive (tonic excitatory tone plus rhythmic, interneuron-selective inhibition) is technically difficult in acute slices, which are largely quiescent without exogenous stimulation. We therefore view CaMKIIα and Thy1 as complementary approximations, each isolating a different circuit interaction. If forced to choose, we’d argue that the CaMKIIα is a better model of disinhibition of excitatory neurons. We will revise the Discussion regarding this point.

      On reproducibility of the voltage imaging findings: We thank the reviewer for this comment and agree that clarification is warranted.

      The voltage imaging dataset combines two levels of analysis with different sample sizes. The population-level firing and spike-correlation analyses (Fig. 5F–H) are pooled across multiple imaging sessions (n = 240 neurons). The spatial clustering analysis of subthreshold voltage correlations (Fig. 6, and the corresponding example traces in Fig. 5A–E) are drawn from a single representative recording session, as the reviewer correctly notes. We have voltage imaging data from 14 fields of view (1 FOV per slice) across 6 mice (240 neurons total; 3–41 neurons per FOV). In revision, we will extend the clustering and spatial-correlation analysis from Fig. 6 across sessions to assess whether the reported organization is reproducible, rather than relying on a single example. We will also revise the text to distinguish clearly which analyses are single-session versus pooled.

      On restricting the computational model of excitatory neurons to stellate cells: We modeled stellate cells as the excitatory population because they are the principal cells reciprocally connected to fast-spiking PV+ interneurons (Fuchs et al., 2016), the interneuron class most directly implicated in theta-nested gamma. Pyramidal cells, by contrast, are primarily connected via 5-HT3a-positive interneurons (Fuchs et al., 2016), with the exception of a subset of "intermediate" pyramidal cells that do show reciprocal PV+ connectivity. Our model, which captures the full measured heterogeneity of stellate cell and PV+ interneuron intrinsic properties and their reciprocal connectivity, is, to our knowledge, the most biophysically constrained implementation of this specific microcircuit to date. Incorporating the PV+-connected intermediate pyramidal population is a natural next step. Because this refinement, which requires more experimental data, is nontrivial and beyond the scope of this study, we will note this explicitly as a limitation of the current model in the revised Discussion.

      In vivo comparison (temporal/phase-locking): We agree that grounding our findings in existing in vivo data strengthens the study and will add these comparisons to the revision.

      Our whole-cell recordings reproduce the temporal organization in vivo and provide further insights into cell-type differences between the principal cells. All cell types were strongly phase-locked to theta, while gamma phase-locking declined across successive spikes, with stellate cells decoupling after the first spike and pyramidal cells after the second. This earlier decoupling in stellate cells may contribute to their weaker theta rhythmicity reported in freely moving rats (Ray et al., 2014; Tang et al., 2014). In extracellular recordings from behaving mice, spike-train cross-correlation identifies putative monosynaptic excitatory connections (1–4 ms) from principal cells onto fast-spiking interneurons (Latuske et al., 2015); the excitation-to-inhibition offset we measured is of comparable magnitude, here resolved as a synaptic-current delay in electrophysiologically classified cell types.

      We note that bursting and theta engagement have been assigned inconsistently across in vivo datasets. Bursty cells are preferentially classified as putative stellate by spikepattern classifiers (Latuske et al., 2015), while anatomically identified pyramidal cells are reported as the bursty, theta-rhythmic population in other work (Ebbesen et al., 2016). Because our cell-type assignments are based on subthreshold intrinsic properties (membrane sag, time constant) rather than spike patterning, our phase-locking results are independent of this classification ambiguity.

      In vivo comparison (spatial organization): We agree high-density silicon-probe datasets are the appropriate reference here. To our knowledge, the anatomical distribution of gamma-locked spiking in superficial mEC has not been characterized in vivo. The highest-density available recordings (Gardner et al., 2022) analyze population activity in the decoded state rather than tissue coordinates, do not examine gamma, and are restricted to grid cells. We regard the dissociation we observe between spatially clustered subthreshold input and spatially distributed spiking as a principal advance of the present study, and as a testable prediction for future high-density recordings.

      Ebbesen CL, Reifenstein ET, Tang Q, Burgalossi A, Ray S, Schreiber S, Kempter R, Brecht M. 2016. Cell Type-Specific Differences in Spike Timing and Spike Shape in the Rat Parasubiculum and Superficial Medial Entorhinal Cortex. Cell Reports 16:1005–1015. DOI: https://doi.org/10.1016/j.celrep.2016.06.057

      Fuchs EC, Neitz A, Pinna R, Melzer S, Caputi A, Monyer H. 2016. Local and Distant Input Controlling Excitation in Layer II of the Medial Entorhinal Cortex. Neuron 89:194–208. DOI: https://doi.org/10.1016/j.neuron.2015.11.029

      Gardner RJ, Hermansen E, Pachitariu M, Burak Y, Baas NA, Dunn BA, Moser M-B, Moser EI. 2022. Toroidal topology of population activity in grid cells. Nature 602:123–128. DOI: https://doi.org/10.1038/s41586-021-04268-7

      Gonzalez-Sulser A, Parthier D, Candela A, McClure C, Pastoll H, Garden D, Sürmeli G, Nolan MF. 2014. Gabaergic projections from the medial septum selectively inhibit interneurons in the medial entorhinal cortex. Journal of Neuroscience 34:16739–16743. DOI: https://doi.org/10.1523/JNEUROSCI.1612-14.2014, PMID: 25505326

      Latuske P, Toader O, Allen K. 2015. Interspike Intervals Reveal Functionally Distinct Cell Populations in the Medial Entorhinal Cortex. Journal of Neuroscience 35:10963–10976. DOI: https://doi.org/10.1523/JNEUROSCI.0276-15.2015

      Ray S, Naumann R, Burgalossi A, Tang Q, Schmidt H, Brecht M. 2014. Grid-Layout and Theta-Modulation of Layer 2 Pyramidal Neurons in Medial Entorhinal Cortex. Science 343:891–896. DOI: https://doi.org/10.1126/science.1243028

      Tang Q, Burgalossi A, Ebbesen CL, Ray S, Naumann R, Schmidt H, Spicher D, Brecht M. 2014. Pyramidal and Stellate Cell Specificity of Grid and Border Representations in Layer 2 of Medial Entorhinal Cortex. Neuron 84:1191–1197. DOI: https://doi.org/10.1016/j.neuron.2014.11.009

      Williams B, Vedururu Srinivas A, Baravalle R, Fernandez FR, Canavier CC, White JohnA. 2026. Fast spiking interneurons autonomously generate fast gamma oscillations in the medial entorhinal cortex with excitation strength tuning ING–PING transitions. eneuro ENEURO.0452-25.2026. DOI: https://doi.org/10.1523/ENEURO.0452-25.2026

    1. eLife Assessment

      This study makes a solid and valuable contribution to elucidating the intricate relationship between mitochondrial calcium and neuronal survival. Well-controlled experiments show that homeostatic mitochondrial calcium correlates with the most resilient neuronal subtypes after optic nerve injury. However, altering mitochondrial calcium levels does not affect neuronal survival as initially predicted by this correlation.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      (2) Some findings can have alternate interpretations that are not considered.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

    4. Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

      We appreciate Reviewer #1’s assessment of our manuscript. We also agree that we should have more clearly indicated that our mitochondrial Ca2+ sensor (Cox8-Twitch2b) is localized to the mitochondrial matrix. The Cox8-mitochondrial localization peptide is a well-established tool first identified in 1992 by Rizzuto and colleagues (Rizzuto, Simpson and Pozzan, 1992). We should have cited this work in our manuscript and will add it to our references. Further, as discussed in our submission, Cox8-Twitch2b has previously been validated for mitochondrial Ca2+ measurements in CNS axons (Witte et al., 2019). Thus, given the decades of use and characterization for this toolset, and the fact that we have pharmacological data supporting mitochondrial matrix localization of Cox8-Twitch2b, we do not feel it is strongly necessary to demonstrate mitochondrial matrix versus inner membrane space localization. However, we could attempt immuno-electron microscopy if this is deemed critical.

      We also agree that the mechanism by which reducing mitochondrial Ca2+ is protective would be satisfying and strengthen this study. But we feel it is beyond the scope of this project. It is likely manifold since mitochondrial Ca2+ impacts many vital cellular functions relevant to pathology including metabolism and apoptosis. We ultimately believe that an adequate investigation of these mechanisms would significantly slow down the dissemination of the core novel findings presented herein.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      We agree with Reviewer #2 that our AAV manipulations of shMCU and MCU overexpression should be analyzed to verify how they alter mitochondrial Ca2+. To do this, we will co-express gene therapy vectors to lower and raise MCU expression with mito-Twitch2b biosensor and perform direct measurements of mitochondrial Ca2+. We will then determine if there is a relationship between gene expression level (inferred by mCherry intensity) and mitochondrial Ca2+ within samples, and if mean mitochondrial Ca2+ levels in treatments are higher or lower than mCherry reporter only controls.

      (2) Some findings can have alternate interpretations that are not considered.

      We will expand our Results and Discussion sections to broaden the interpretations of our data.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

      We agree that a more fine-grained understanding of RGC mitochondrial Ca2+ diversity would make interpretations of our data stronger. In our revisions, we will thus expand the number of RGC families in which we directly measure homeostatic mitochondrial Ca2+ levels. To do this, we will perform in vivo mito-Twitch2b measurements, collect and fix retinal wholemounts and immunostain for ON-OFF-direction selective RGCs using the marker CART and F-RGCs using the marker Foxp2. This will provide a complement of well-surviving RGC types (alpha and intrinsically photosensitive RGCs already examined) and poorly-surviving types.

      Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

      We agree with the feedback from Reviewer #3, especially as it aligns with input from other reviewers. As these points agree with aspects above we will briefly reiterate our proposed revisions. We will validate the true effects on mitochondrial Ca2+ levels after gene therapy treatments by co-injecting AAV-mito-Twitch2b and AAV-shMCU or AAV-MCU. We will measure mitochondrial Ca2+ levels and correlate these levels with mCherry reporter expression intensity to determine the effect size of these treatments, and compare sample mean mitochondrial Ca2+ levels with those of mCherry control AAV.

      To further map the variance in homeostatic mitochondrial Ca2+ levels to RGC types we will perform in vivo mito-Twitch2b imaging, and then immunostain for ON-OFF-direction selective RGCs (CART) and F-RGCs (Foxp2), two poorly surviving RGC types.

      Lastly, we agree with Reviewer #3 that finer delineation between mitochondrial Ca2+ levels and their relationship to survival may be informative. We will split RGCs into smaller subgroups based on homeostatic mitochondrial Ca2+ levels and examine their survival outcome.

      Overall, we thank the Reviewers for their feedback, and believe the suggested changes will greatly strengthen our study.

      REFERENCES

      Rizzuto R., Simpson A.W. and Pozzan T. (1992). Rapid changes of mitochondrial Ca2+ revealed by specifically targeted recombinant aequorin. Nature, 358 (6384): 325-327.

      Witte M.E., Schumacher A-M., Mahler C.F., Bewersdorf J.P., Lehmitz J., Scheiter A., Sanchez P., Williams P.R., Griesbeck O., Naumann R., Misgeld T. and Kerschensteiner M. (2019). Calcium influx through plasma-membrane nanoruptures drives axon degeneration in a model of multiple sclerosis. Neuron, 101(4): 615-624.

    1. eLife Assessment

      This article describes the comprehensive metabolic phenotype of a mouse model of Down Syndrome, together with supporting transcriptomic, metabolomic, and biochemical data. The evidence presented is compelling and highlights several core phenotypes including insulin resistance, dyslipidemia, and tissue signatures indicating inflammatory and cellular stress pathways. Similarities and differences in male and female mice are highlighted. This important study provides essential groundwork for the further genetic dissection of dosage-sensitive genes causing metabolic dysregulation in Down Syndrome.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

    3. Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

    4. Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

      Weaknesses:

      Many aspects of this study have been described in other Down syndrome mouse models, though there are certainly aspects that are new. It would be useful if the authors could do a direct critique and comparison with previous publications in the area, utilizing the same Down Syndrome mouse model. There are also a few limitations in the number of animals used and the interpretation of the data that should be acknowledged.

      We have cited all relevant publications using Down syndrome mouse models. Regarding the Dp16 model, we have cited and discussed the only other study addressing metabolic aspects beyond body weight (Reference #138; PMID: 39803786). While that study reported glucose intolerance, insulin resistance, and defective insulin secretion, we did not measure pancreatic insulin content in our mice. Crucially, while the previous study found no sexual dimorphism, our study observed extensive sexual dimorphism in body weight gain, tissue-specific gene expression, and serum and liver metabolite changes.

      Regarding sample size, we used 6 mice per genotype per sex for transcriptomic and metabolomic analyses; this is constrained by the cost of performing these omics-type analyses. For mitochondrial respiration assays, we used 9–10 mice, and for most other in vivo and ex vivo assays, we utilized 12–15 mice, with some assays exceeding 20. We believe these sample sizes are robust and appropriate for this study.

      Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

      Weaknesses:

      (1) Assessing metabolism using dynamic testing is a strength. ITT, GTT and LTTs are included.

      (2) The dosing for GTTs, ITTs and LTTs was performed per body weight. But the mice under chow and HFD had different body weights. This may compromise the interpretation of the data. Further, ITTs are presented as percentage change, and this can be heavily influenced by baseline glucose measures. The changes appear quite dramatic, so can the authors plot the raw data instead?

      We have updated the ITT data plots to show raw glucose values instead of percentage change. Regarding the dosing, we believe basing it on body weight is an appropriate approach. This method is consistent with nearly all published rodent studies, as blood volume and metabolic tissues such as skeletal muscle and adipose tissue scale with body weight. Adjusting for weight prevents potentially erroneous conclusions. As for the diet groups, we compared WT and Dp16 mice only within the same diet group (Chow or HFD) rather than across different diets. We believe this ensures a valid and appropriate comparison for our study.

      (3) In addition, throughout the manuscript, it is not clear which tissues are the most dominant in disrupting metabolism. The ITT and GTT are composite measures across tissues. Tissue-specific analyses using a clamp technique or isolated tissues may provide more clarity here.

      Our data suggest a systemic metabolic deficit across multiple tissues, supported by tolerance tests, pan-tissue transcriptomic analyses, and liver and serum metabolite profiling. This is consistent with the triplication of genes in Down syndrome, several of which have known metabolic roles as highlighted in our discussion. We do not have evidence to support the role of a dominant tissue that contributes to the systemic metabolic dysfunction.

      Regarding the suggestion to use a clamp technique, we agree this would effectively determine whether insulin resistance is localized in the liver or skeletal muscle. However, we do not currently have the necessary equipment at Johns Hopkins University to perform these experiments. Conducting this work would require sending separate cohorts of WT and Dp16 male and female mice (on both chow and HFD) to an NIH-funded Mouse Metabolic Phenotyping Centre (MMPC). While we appreciate the value of this approach, we believe such labor-intensive experimentation falls beyond the scope of the present study.

      (4) One of the aims of the study was "to help inform the selection of mouse model that best reflects the metabolic profile of DS". The discussion does not contain a comparison between the previous work on different strains and relative to known human data.

      We chose not to include a comparison of different mouse models in the "Discussion" section because we previously highlighted the widely used Down syndrome models (Ts65Dn, Tc1, and TcMAC21) and their associated caveats in the "Introduction." Given the significant limitations of those models such as hypermetabolism in TcMAC21 and the presence of 41 triplicated protein-coding genes unrelated to human chromosome 21 we focused our in-depth metabolic analyses on the Dp16 model, which does not share these issues. We felt that restating this information in the "Discussion" would be unnecessarily repetitive.

      (5) Data availability. Raw metabolomic data should be made available.

      We have uploaded all metabolomics data, along with details regarding sample processing and data analysis, to the Metabolomics Workbench, an NIH-funded public repository. We have updated the "Methods" and "Data Availability" sections of the manuscript to include this information and the corresponding access link.

      Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

      Weaknesses:

      It might have been helpful to have included blood pressure measurements, given the differences in 19-Nor-deoxycorticosterone. The discussion references several articles that describe sex-dependent differences in metabolic phenotypes in humans with Down syndrome, and it might have been helpful to state more explicitly whether these differences correlate with those observed here in mice.

      We appreciate the suggestion of blood pressure measurements. While we agree this is an important metric, given the metabolic focus of the present study and the significant volume of data already presented, we feel that blood pressure analysis is beyond the current scope and better suited for a follow-up study.

      Our study highlights sex differences in metabolic phenotypes in individuals with Down syndrome. While most published human studies focus on a limited set of parameters such as body weight, adiposity, serum lipoprotein profile, and fasting lipid/glucose levels our mouse data remain generally concordant with these findings. Beyond these standard measurements, we also observed substantial sex differences in pan-tissue transcriptomes as well as serum and liver metabolites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A major question is how these findings compare to data that have previously been published. For example, Lamantia et al. Bone 2024 and Dard et al. European Journal of Pharmacology 2025 both report no changes in body weight using the same Dp(16)1Yey Down syndrome mouse model? There is also a recent publication on liver dysfunction in Down Syndrome using the same mouse model. It would be useful to understand some of the similarities and differences of what is being reported by Dunn et al. Cell Rep 2026. In this assessment, there is an in-serum alanine transaminase (ALT) level, which was not the case in Dunn et al?

      For the Lamantia et al. Bone 2024 study, the authors only measured the body weights of Dp16 mice at 6 weeks of age. Our findings at 6 weeks align with Lamantia et al., showing no weight differences between Dp16 and WT mice of either sex (Fig. 2A and C). For the Dard et al. 2025 study, the authors only measured the body weights of Dp16 mice at 12 weeks old (P90) and observed no differences in body weights between genotype of either sex. At 12 weeks of age, we also did not observe body weight differences between Dp16 male mice and WT littermates (Fig. 2A). However, at 12 weeks of age, the Dp16 female mice clearly gained more weight compared to WT littermates (Fig. 2C). Our study tracked weights weekly from 6 to 16 weeks, revealing that while Dp16 females start at weights similar to WT littermates, the groups diverge over time. The reason for the difference between our findings and the single-point measurement by Dard et al. is unclear. Notable variables include:

      Mouse Sourcing: We obtained all cohorts and littermate controls from Jackson Laboratory, while Dard et al. bred their mice in-house.

      Diet: We used Envigo standard chow (catalogue # 2018SX). Dard et al. did not specify the chow used in their study.

      It remains uncertain whether these or other environmental factors contribute to the observed weight differences in female mice.

      In the Dunn et al study (Cell Rep 2026), they also performed metabolic analyses on serum and liver tissue in Dp16 mice. Consistent with their metabolic analyses of serum and liver tissue in Dp16 mice, we also observed the upregulation of multiple bile acids, including taurochenodeoxycholic, tauromuricholic, taurolithocholic, and lithocholic acids. Furthermore, our findings align with theirs regarding the transcriptomic and biochemical signatures of hepatic inflammation and fibrosis. However, there are two notable differences between our studies:

      (1) Liver Injury Markers: We observed an elevation in serum ALT, whereas the Dunn et al. study did not.

      (2) Sex Differences: We identified significant sex differences in the Dp16 transcriptome and metabolome. In contrast, Dunn et al. reported minimal to no sex differences and consequently combined male and female data for all analyses.

      Because Dunn et al. combined male and female data, a sex-stratified comparison between our results (separated by sex) and theirs was not feasible.

      (2) It would be important to understand trends in wild-type animals compared to Dp16 mice. For example, the sex specific and non-specific features - are any of these described in obesogenic wild-type animals fed on a high-fat diet? I.e., are the same features at play and just exacerbated in Dp16, or is this a Dp16-specific feature of systemic metabolism?

      Published literature indicates that WT females typically gain significantly less weight on a high-fat diet (HFD) than WT males. However, our data suggest that the weight gain patterns observed in Figure 6A and C are specific to the Dp16 genotype. Dp16 females gained substantially more weight during the first six weeks of HFD before WT females caught up. In contrast, Dp16 males showed robust initial weight gain comparable to WT controls, but their weight plateaued after seven weeks while WT controls continued to gain, leading to a clear divergence (Fig. 6A).

      Other metabolic parameters also appear specific to the Dp16 model. On a standard chow diet, WT mice of both sexes generally do not exhibit glucose intolerance, insulin resistance, dysregulated lipoprotein profiles (VLDL-TG), or an impaired capacity to handle lipid loads. We observed all of these features in our Dp16 male and female mice (Fig. 3). Furthermore, transcriptomic analyses of Dp16 mice on standard chow revealed gene signatures of inflammation, fibrosis, and oxidative stress that are absent in WT mice.

      When challenged with HFD, while WT mice typically develop glucose intolerance and insulin resistance, the triplicated genes in Dp16 mice significantly exacerbated this metabolic deterioration. This is reflected in the worsening of glucose control and insulin sensitivity observed in our tolerance tests.

      In summary, most of these metabolic features are specific to Dp16 mice on a standard chow diet and are further exacerbated when combined with a high-fat diet.

      (3) Food intake data is difficult to interpret when weight has already diverged, as bigger animals will eat more food. Hence, the higher food may be a consequence rather than a cause of the weight gain (data in Figure 1).

      The reviewer makes a valid point. Since physical activity and energy expenditure do not differ significantly between Dp16 females and WT controls (Fig. 2F), the observed increase in food intake may indeed contribute to the higher body weights in Dp16 female mice.

      To rigorously confirm this, food intake would need to be measured between 6 and 8 weeks of age, prior to the divergence in body weight. Unfortunately, we did not measure food intake at that earlier time point.

      (4) The n numbers seem to vary significantly. For example, the use of n=6 for metabolic studies is generally rather small and underpowered. For the seahorse data, another concern is the snap freezing of samples before Seahorse assessment. For example, snap freezing of samples has been shown to increase certain metabolites. Freeze-thaw tissues often show a significant reduction in optical redox ratio.

      Regarding the transcriptomics and metabolomics studies, we utilized N=6 mice per tissue per sex. While we agree that a larger sample size is always preferable, the high cost of OMICS analyses covering 144 RNA-seq and 48 metabolomics samples limited our capacity to increase this number. However, N=6 remains a robust and standard approach for these specific assays. For the majority of our other in vivo and ex vivo data, we employed a higher sample size of 12-15 mice per genotype per sex to ensure statistical rigour. For a few assays, we have sample size of over 20.

      Regarding the respirometry analysis, we acknowledge the limitations of using frozen tissue. We chose this method because it allowed us to perform Seahorse assays on multiple tissues from 9-10 mice, which is a significant sample size for this type of analysis. The alternative isolating mitochondria from fresh tissue would have restricted our ability to process multiple tissues from a large number of animals on the same day due to the length of the protocol. We believe this trade-off was necessary to maintain a high sample size across various tissues.

      (5) For oestradiol measurements, were the samples taken at the same times within the estrous cycle? This may affect the comparability of female Dp16 and WT mice?

      Regarding our protocol, blood samples were collected between 11:00 AM and noon, with food removed two hours prior. While we did not specifically monitor the oestrous cycle of the female mice, serum samples for both the Dp16 females and WT littermates were collected on the same day and at the same time to ensure comparability across the groups.

      (6) Body weight reduction and organ size reduction on an HFD are especially interesting. Could enhanced inflammation and fibrosis be the root cause of this? Are there other mouse models where this is the reason?

      On a high-fat diet, we observed a reduction in iWAT and gWAT fat depot weights in both male and female Dp16 mice, which is consistent with their lower overall body weights (Fig. 6 - figure supplement 3). Conversely, Dp16 females fed a high-fat diet showed increased heart and kidney weights. Despite their lower adiposity, the Dp16 mice on this diet exhibited greater insulin resistance and glucose intolerance (Fig. 7). This suggests that the worsening of glucose control is independent of obesity. While we observed signatures of inflammation and fibrosis, we do not yet have direct mechanistic evidence demonstrating that these factors causally impaired glucose and lipid metabolism.

      (7) The authors are circumspect throughout to avoid over-claiming, as the majority of data is observational. One exception: "Many bile acids serve as ligands for nuclear hormone receptors (e.g., FRX and TGR5) that control various aspects of glucose and lipid metabolism (74, 75), and extensive changes in circulating bile acids are contributing, at least in part, to the systemic metabolic phenotypes in Dp16 mice." The authors have not shown a direct link between bile acids and metabolism in this model. Please edit.

      We have edited the text accordingly.

      Minor:

      (1)"Most human studies at the whole-body level are limited to assessing the impact of trisomy 21 on food intake, adiposity, physical activity level, and energy expenditure in adolescents or adults with DS"

      While we were uncertain of the reviewer's specific intent regarding the suggested changes, we have rephrased the sentence for clarity.

      (2) It is somewhat surprising that T3 is elevated, although there are reports of T3 elevation in visceral obesity in humans (e.g., Sun Nam et al., Obes Res Clin Pract, 2010).

      We observed that T3 levels did not differ by genotype in mice of either sex when fed a standard chow (Fig. 2 - figure supplement 5). However, we noted elevated T3 levels in both male and female Dp16 mice on a high-fat diet (Fig. 6 - figure supplement 2). While increased T3 levels correlated with higher physical activity and a modest increase in metabolic rate in Dp16 females, this was not observed in males (Fig. 6). We do not currently have a clear explanation for these findings. Given that individuals with Down syndrome often present with hypothyroidism and lower T3 levels, this discrepancy may reflect a species-specific difference between humans and mice.

      (3) Please can the authors clarify the percentage gene coverage, as this is quoted as ~58% of Hsa21 gene orthologs or ~65% of the Hsa21 gene orthologs, where the same reference is used.

      We apologize for the confusion. The number of triplicated genes in Dp16 mice corresponds to ~58% of Hsa21 genes (PMID: 26765563). We have corrected the typographical error in the text.

      (4) "segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs" for this given percentage majority sounds too strong, and the use of percentage is recommended.

      We have modified the text accordingly.

      (5) It is puzzling that in female gWAT with 7 triplicated Hsa21 gene orthologs (Rbm11, Chodl, Cldn8, Sh3bgr, Igsf5, Itgb2l, and Tmprss2). Could this be a technical issue? Was the reduced expression quantified by RT-Q-PCR?

      We have examined the normalized counts in the RNA-seq data for the seven genes in question, and the results do not appear to be an artifact. The sample size for this data is six mice per tissue per sex. In general, we prefer utilizing raw and normalized counts from RNA sequencing because there is a linear relationship between transcript amount and raw counts that is independent of housekeeping genes. In contrast, RT-qPCR involves mRNA amplification and requires expression to be normalized by one or more housekeeping genes (such as GAPDH, β-actin, 36B4, or ubiquitin) under the assumption that their levels remain constant.

      (6) The difference in body temperature is of interest. In male Dp16 mice, there is an increase in core temperature and a lowering of body temperature in females. In female Dp16 mice, higher estradiol levels have been stated by the authors to contribute to lower body temperature and higher physical activity (69-72). I am uncertain if the references are all relevant, as some relate to ovariectomized animals. No explanation is given for males.

      We currently do not have an explanation for why Dp16 males on a chow diet exhibit higher core body temperature, while Dp16 females show lower body temperatures. Although elevated T3 levels can increase body temperature, we have ruled this out; our data indicates there are no significant differences in T3 levels between genotypes for either sex on a chow diet.

      (7) The authors find a higher percentage heart weight in Dp16 mice on HFD and comment in the discussion that this is in keeping with "high-fat diet-induced cardiac hypertrophy". From what I can see, no histology has been performed to justify this statement. Furthermore, it would be useful to understand which animals had congenital heart disease in the first instance.

      We have modified the text accordingly. Unfortunately, we do not have histology data on the heart to inform us on whether some of our mice had congenital heart disease.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should comment on the dosing method of glucose/insulin/lipid in the tolerance tests to acknowledge that differences in body weight may affect these tests. In addition, I encourage the authors to present ITT data as raw data, and not % change.

      In response to the reviewer’s comments, we have updated the ITT data plots to show raw data rather than percentage change. Regarding the dosing methodology, we maintain that basing dosage on body weight is appropriate. This approach is consistent with the vast majority of published rodent studies, as blood volume and metabolic tissues—such as skeletal muscle and adipose tissue—scale with body weight. Standardizing dose independently of body weight could lead to erroneous conclusions.

      (2) It would be useful for the authors to include a discussion on the likely specific tissue involvement in the whole-body metabolic disturbance. From my reading of the manuscript, there seems to be data suggesting functional and transcriptional dysfunction across most tissues, but do the authors suggest there is a dominant tissue in this regard?

      Due to the triplication of large number of genes on human chromosome 21, people with Down syndrome exhibit deficits across most organ systems (PMID: 32029743). Metabolic homeostasis also involves multiple tissues and cell types (adipose tissues, liver, skeletal muscle, pancreas, gut, hypothalamus, and immune cells). Most of the triplicated genes do express across these tissues. Our data indicate metabolic dysregulation across adipose tissues (white and brown), liver, skeletal muscle, and hypothalamus. Given the complex genetic perturbations of the Down syndrome mouse model, we do not think that there is a dominant tissue that contributes disproportionately to the systemic metabolic dysfunction phenotypes we observed in the Dp16 mice. Rather, we think that the metabolic phenotype is due to the combined deficits across multiple organs and tissues. As we do not have data to support the disproportionate contribution of any one tissue, we therefore did not speculate on the dominant contribution of any single tissue in the Discussion.

      (3) Related to this, muscle lipid is thought to be a major driver of muscle insulin resistance. Do the authors have measures of muscle lipid accumulation? This might be particularly interesting in the HFD models.

      Unfortunately, we did not measure lipid content in the skeletal muscle during this study. For the chow-fed mice, the entire gastrocnemius muscle was used for RNA isolation to perform RNA sequencing, and no tissue remains for additional analysis. Regarding the HFD-fed group, skeletal muscle was not collected at the termination of the study. As a result, we are unable to provide the requested lipid analysis data.

      (4) For mitochondrial analyses - do the authors have measures of total tissue mitochondria, and might changes in mitochondria abundance be driving some of these differences?

      For all our mitochondrial respiration analyses, we normalized the data to mitochondrial content as quantified by the MTDR assay (PMID: 32432379; PMID: 39704485). These results indicate that for a given amount of mitochondrial content, respiration as measured by the Seahorse assay is reduced in Dp16 mouse tissues, specifically in the BAT and liver.

      (5) To broaden the scope and interest, can the authors compare the transcriptional or metabolomic data to what has been found in non-DS insulin resistance (humans or mice), for example? This may help to highlight the key changes in metabolism that are causal for specific phenotypes.

      Overall, this is a comprehensive assessment of metabolism in a DS model.

      We appreciate the reviewer’s suggestion. However, given the vast number of published datasets on non-DS insulin resistance in both humans and mice, comparisons would yield varying results depending on the specific datasets selected. Consequently, we feel that such an analysis is beyond the scope of this study. We would like to highlight that many of the processes dysregulated in Dp16 mice as identified through our pan-tissue transcriptomes and metabolomes align with those frequently observed in non-DS insulin resistance. These include signatures of chronic low-grade inflammation, fibrosis, ER and oxidative stress, and impaired glucose and lipid metabolism.

      Reviewer #3 (Recommendations for the authors):

      It is slightly disconcerting that Figure 5 - Figure Supplements 2-5 are referred to in the text before the data in Figure 5 are discussed. It might make sense to indicate that the data are discussed further below (assuming that the authors do not wish to renumber these figures).

      We have fixed this issue raised by the reviewer.

    1. eLife Assessment

      This study presents an important large-scale behavioral and transcriptomic analysis of Drosophila that are heterozygous for putative loss-of-function alleles of homologs of human genes that have been linked to autism spectrum disorders. The authors consider 48 genes as hits from their screen, which show significant behavioral alterations in sleep, basal activity, and/or social behavior, and significant sexual dimorphism. The authors then focus on the domino/SRCAP gene as a candidate regulator of sleep, social behavior, transcriptional programs, and RNA splicing. The work generates a solid dataset and applies quantitative analytical approaches that will be of interest to researchers in the field, yet the evidence presented remains incomplete because issues of genetic background need to be further addressed.

    2. Joint public review:

      Summary:

      In this study, Stirtz et al., performed a targeted screen of 80 Drosophila strains carrying heterozygous MiMIC insertions in genes that are homologous to human genes that have been linked to autism spectrum disorders (ASD). This is an important and timely topic, as human genetic studies have identified a large number of ASD risk genes, yet the functional characterization of many of these candidates remains limited. The authors identify 48 putative mutants with altered sleep, activity, or social behavior. They then focus on one hit, domino (the orthologue of human SRCAP), for which the heterozygous MiMIC mutants show altered behavior in males but not in females. They show that domino is a candidate regulator of sleep, activity, social behavior, transcriptional programs, and RNA splicing. The authors molecularly validate that the heterozygous MiMIC insertion in domino causes a 50% reduction in gene expression, and use RNA-seq to show that the heterozygous MiMIC males and females have altered gene expression profiles and splicing patterns. Finally, they use immunostaining against the commonly used synaptic marker, Bruchpilot, to show that both males and female heterozygous domino flies express a higher immunosignal compared to the wild-type control.

      Strengths:

      This work provides potential genetic links between human ASD genes and fly behavioral phenotypes. Overall, it represents an ambitious and technically valuable effort that generates a substantial behavioral dataset across a large number of ASD-associated orthologues and develops quantitative analytical approaches to extract information from complex phenotypes. One strength of this study is its focus on heterozygous mutants, which is more representative of human scenarios. The study also provides a potentially useful resource for the field, particularly through the identification of candidate genes and behavioral signatures that may warrant future mechanistic investigations. The screening experiments and analysis are well conceived, the manuscript is very clearly written and is easily understandable, and the concise, accurate interpretations for each result, aided by clear graphic representation of multiple dimensions in the behaviors tested, allow the reader to understand the paper with ease.

      Weaknesses:

      The work presents a few important weaknesses, especially with regard to the genetic and molecular validation of the mutants identified.

      (1) The authors validate that the MiMIC insertion affects the gene of interest only for the domino gene. The original MiMIC study (PMID: 25824290, eLife) reported that ~8% (5/63) MiMIC lines do not function as strong loss-of-function alleles. Thus, of the 48 hits identified here, one would estimate that ~4 of them may not cause the loss of function of the gene defined by the MiMIC insertion. To strengthen their claim, the authors would need to confirm that all of the MiMIC lines that they consider as hits do indeed significantly reduce the expression of the target genes.

      (2) Although the authors document that they validated the phenotype seen in the domino MiMIC line using a second mutant allele (Trojan), these two mutants share the same genetic background because the Trojan line was made from the MiMIC line via recombinase-mediated cassette exchange. Thus, the phenotype seen in the MiMIC and Trojan lines would need to be confirmed using a completely independent mutant in order to demonstrate that the reported behavioral, molecular, and synaptic defects reported can be fully attributed to the partial loss of domino function. Also, while the authors performed an RNA-seq experiment in both the MiMIC and Trojan lines, they do not show whether the Bruchpilot phenotype is also seen in the Trojan allele. Thus, this phenotype would also need to be examined in the Trojan allele or, preferably, in a mutant allele that is independent of the MiMIC line.

      (3) The RNA-seq results would benefit from a discussion of potential compensatory or secondary transcriptional effects resulting from the constitutive domino reduction, particularly since the expected global bias toward transcriptional downregulation was not observed. In addition, some neurobiological interpretations appear stronger than currently justified by the literature or the data presented, particularly regarding the Bruchpilot immunoreactivity analyses and their relationship to sleep-regulatory circuits. Additional validation using better-established sleep-related neuronal populations, together with a clearer discussion of sex-specific effects and alternative interpretations of the observed phenotypes, would substantially strengthen the manuscript.

      (4) An explanation of the extensive PCA analyses performed would help the naïve reader.

    1. eLife Assessment

      In this useful Tools & Resources article, the authors describe a new cryogenic light microscopy design and characterize its temperature and spatial stability. This compelling system avoids the challenges associated with vacuum-based designs, particularly vacuum transfer systems, which are difficult to engineer. A key advantage of the system is that it reduces ice contamination and drift, which are the primary challenges in open cryostat systems.

    2. Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      We thank the reviewer for their positive assessment and strive to address the weaknesses they have constructively raised below.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we will include a new supplementary figure (Fig. S2) providing detailed 3D views of the copper adapter, microscope objective, and plastic cap. The figure will show that the rim surrounding the front lens of the objective is covered by the plastic cap to provide thermal insulation between the objective housing and the cryogenic environment (Fig. S2b). We will also clarify that the front surface of the cap is levelled with the front objective lens to maintain the full working distance of the objective while allowing for axial movement of the z-stage. Finally, we will explicitly state that the front objective lens is exposed to the cryogenic environment (cold nitrogen gas).

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      To minimise vibrations arising from the nitrogen refilling pumps, the cryostat and liquid nitrogen tubing are mechanically decoupled from the microscope cage system, objective, translation stages, and sample. Specifically, the cryostat and nitrogen tubing are supported independently on a laboratory jack and surround the cage system without rigid mechanical contact. In the revised manuscript, we will update Fig. S1a, b to illustrate the liquid nitrogen tubing and refilling system more clearly. In addition, we will include a new supplementary figure (Fig. S3) to show the detailed cryostat design, refilling tubing, and temperature sensor position.

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      We agree that the purpose of the rectangular aperture should be made clearer. In line with the suggestion from the reviewer, in the revised manuscript, we will briefly explain that the aperture is intended for selective illumination, such as in applications to cryo-FIB lamellae, and will cite our recent preprint describing this approach.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      We thank the reviewer for this suggestion. In the revised manuscript, we will include an additional supplementary figure (Fig. S4) showing the axial drift measured without focus-lock correction to illustrate the intrinsic mechanical stability of the microscope. We will also clarify that the periodic lateral displacement observed along the x-direction (approximately 300 nm amplitude with a period of ~22 minutes) arises from slight lateral repositioning accompanying z-stage stepping during focus-lock operation, likely due to mechanical coupling between the axes of the translation stage. We will revise the lower panel of Fig. 2d by reducing the y-axis range to improve data visibility.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      We agree that this would be useful. In the revised manuscript, we will update Fig. 3 to include two additional low and intermediate-magnification cryo-EM images showing a representative grid square and a zoomed-in region of it, including a few grid holes. These images provide an overview of the ice contamination across a substantially larger field of view.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      The y-axis in Fig. 4b represents the localization rate (number of detected localizations per frame) within the selected ROI in Fig.4c and was not normalized. The values were calculated in SMAP by binning the localization frames into 100 temporal bins and dividing the number of localizations in each bin by the corresponding bin width, resulting in units of localizations per frame. Therefore, values close to 1 indicate approximately one localization detected per frame at that time point. To avoid potential confusion regarding the interpretation of this representation, we will replace this plot in the revised manuscript with a more explicit visualization showing the number of detected localizations per defined number of frames as a function of time (frame number) for the specific ROI shown in Fig. 4c.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

      In the revised manuscript, we will add fluorescence intensity traces corresponding to the single-molecule events shown in Fig. 4c and Fig. 4f to further demonstrate their single-molecule emission characteristics.

      Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      We thank the reviewer for their critical assessment, and for their suggestions below which we have used to improve the manuscript.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      We thank the reviewer for highlighting this point. We agree that focus stabilization strategies for cryo-FIB lamellae are likely to differ from those used for the individual bacterial cell samples. For lateral drift correction, the presence of a single continuously detectable bright feature within the field of view is sufficient. Importantly, this feature does not need to be a fluorescent bead; any stable signal that can be continuously detected by the camera can serve as a suitable reference for drift correction. We will expand the Discussion to describe potential strategies for stable cryo-SMLM imaging, including the use of intrinsic sample or lamella features for autofocus, minimal fiducial-based approaches, and the practical implications of dedicating the far-red channel to focus stabilization.

      Furthermore, in the revised manuscript, we will include a new supplementary figure (Fig. S4) demonstrating the intrinsic axial stability of the microscope in the absence of active focus-lock correction. These measurements show that the system remains within the objective's depth of focus for a relatively long time, providing adequate stability for experiments in which far-red fluorescent fiducial beads are unavailable, such as cryo-FIB lamella imaging.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      We thank the reviewer for this thoughtful comment. We agree that a correlative example displaying more pronounced structural features would further illustrate the capabilities of cryo-SMLM. However, the primary aim of the present work is the development and characterization of a robust cryogenic super-resolution microscope for reliable cryo-SMLM and correlative cryo-CLEM, rather than the demonstration of new biological applications. The utility of correlative cryo-SMLM/cryo-ET for resolving cellular structures has already been established in previous studies, including those employing rsEGFP2-labelled targets.

      The correlative dataset presented here is intended to demonstrate the compatibility of the microscope with cryo-CLEM workflows rather than to provide detailed biological insight. Moreover, the use of intact E. coli cells imposes inherent limitations on the ultrastructural information accessible by cryo-electron tomography; overcoming these limitations would typically require specimen thinning, for example, by cryo-focused ion beam (cryo-FIB) milling, which is beyond the scope of the present work.

      Regarding the concern about auto-fluorescence, elevated background fluorescence is not unique to bacterial samples or green fluorescent proteins but is a general consideration in cryo-SMLM that depends on the specimen and imaging conditions. While auto-fluorescence may reduce image contrast, it does not affect the conclusions of this work, which focuses on the design and performance of the microscope.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

      Thank you for this useful suggestion. In the revised manuscript, we will make available the complete SolidWorks CAD files for all custom-designed components, together with a comprehensive parts list and the full assembly corresponding to Fig. S1 as supplementary materials.

    1. eLife Assessment

      This manuscript describes a valuable study of the mechanism by which acetylation on the histone H3 core domain regulates RNA polymerase II transcription passing through nucleosomes. The authors provide convincing evidence that acetylation influences transcription in a context-specific fashion. Some questions relating to the static nucleosome structures and the polymerase passage remain, but this manuscript will be of considerable interest to researchers in the chromatin and transcription fields.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate how site-specific acetylation within the histone H3 folded domain affects RNA polymerase II transcription through nucleosomes. They focus on H3K56ac, H3K64ac, and H3K122ac, prepare chemically defined nucleosomes carrying each modification, and compare their effects using an in vitro transcription assay, cryo-electron microscopy structures, and micrococcal nuclease sensitivity assays.

      The main finding is that H3K56ac and H3K122ac increase production of full-length run-off transcripts and reduce pausing near the nucleosomal dyad region, whereas H3K64ac has little detectable effect under the same reconstituted conditions. The structural analyses suggest that H3K56ac weakens or destabilizes DNA near the entry/exit region, while H3K122ac alters histone-DNA contacts near the dyad. These observations support a model in which different acetylation sites within the H3 folded domain influence nucleosomal transcription barriers through distinct local effects on histone-DNA interactions.

      This is a useful study because it examines histone core-domain acetylation using chemically defined nucleosomes and directly compares several modifications in the same experimental system. However, the broader cellular context of these modifications is not sufficiently developed, and some mechanistic conclusions rely on correlations between static nucleosome structures and endpoint transcription assays rather than direct observation of polymerase passage through modified nucleosomes.

      Strengths:

      (1) The study uses site-specifically acetylated H3 proteins and reconstituted nucleosomes, allowing direct comparison of H3K56ac, H3K64ac, and H3K122ac under controlled conditions.

      (2) The combination of transcription assays, cryo-electron microscopy, and nuclease sensitivity assays provides multiple lines of evidence, particularly for increased DNA end flexibility in H3K56ac nucleosomes.

      (3) The authors analyze unmodified, H3K56ac, H3K64ac, and H3K122ac nucleosomes in parallel, with reported structural resolutions of approximately 3 Angstroms and accompanying validation materials.

      (4) The negative result for H3K64ac is informative, because it distinguishes the direct effect of this modification in a minimal reconstituted system from prior cellular associations with active chromatin and histone eviction.<br /> The comparison with H3 N-terminal acetylation highlights that acetylation within the folded domain may affect transcription at different positions or by different mechanisms than tail acetylation.

      Weaknesses:

      The rationale for focusing on H3K56ac, H3K64ac, and H3K122ac has not been developed sufficiently. The manuscript would benefit from a clearer summary of what is known about the abundance of these modifications in cells, the enzymes or histone metabolic pathways that may introduce or remove them, and whether they are thought to occur before histone deposition, on assembled nucleosomes, or during nucleosome remodeling.

      The central mechanistic model is based mainly on correlations between structures of free nucleosomes and endpoint transcription assays. The study does not directly observe RNA polymerase II paused at or passing through the relevant nucleosomal positions, so the proposed link between local structural changes and reduced pausing should be stated with appropriate caution.

      The H3K56ac interpretation is supported by both structural observations and nuclease sensitivity data, but the map comparison underlying the reduced entry/exit DNA density is still mostly qualitative. The manuscript should more clearly state the map comparison conditions, such as contouring and local map quality, so that non-specialist readers can judge how robust the local density differences are.

      The H3K122ac mechanism is plausible, but the evidence for dyad destabilization is more indirect. The main support comes from the orientation of the K122 side chain and its distance from DNA, while an independent biochemical test of dyad-region destabilization is not provided.

      The transcription assay appears to include statistical testing, but the figure legend and methods should more clearly state which tests were used, what comparisons were made, how n was defined, and whether multiple-comparison correction was applied.

      The relationship between the 198 bp transcription template, the linker DNA, the 9-base mismatched region, and the DNA regions modeled in the cryo-electron microscopy structures is somewhat difficult to follow. This does not necessarily require new experiments, but a clearer explanation would help readers connect the transcription assay design with the structural models.

      The use of H3.2 C110A for chemical ligation and the use of the PL2-6 single-chain antibody fragment for cryo-electron microscopy sample stabilization are reasonable technical choices, but their purposes and possible effects on interpretation should be explained more clearly for readers outside structural biology.

      Because the work uses a minimal in vitro system with human nucleosomes and Komagataella phaffii RNA polymerase II/TFIIS, the conclusions should be limited to direct physical effects on nucleosome transcription barriers unless cellular cofactors, remodelers, histone chaperones, additional modifications, and nucleosome positioning are addressed or discussed.

    3. Reviewer #2 (Public review):

      Summary:

      Chromatin regulates a wide range of biological processes. The nucleosome, composed of 147 bp of DNA wrapped around a histone octamer containing histones H2A, H2B, H3, and H4, is the fundamental unit of chromatin. Post-translational modifications of histone proteins regulate the dynamic properties of nucleosomes and thereby influence chromatin accessibility and gene expression. Among these modifications, lysine acetylation on histone H3 is closely associated with transcriptional activation. While the epigenetic functions of acetylation on the histone H3 N-terminal tail have been extensively studied, the molecular mechanisms by which acetylation within the histone H3 core domain, particularly at Lys56, Lys64, and Lys122, modulates nucleosome architecture to facilitate RNA polymerase II (RNAPII) transcription remain unclear.

      In this study, Oishi et al. investigated the effects of histone H3 acetylation at K56, K64, and K122 on RNAPII transcription using in vitro transcription assays. Furthermore, the authors determined the three-dimensional structures of nucleosomes containing these acetylation marks by cryo-electron microscopy single-particle analysis, revealing distinct structural dynamics depending on the acetylation site. Overall, this study advances our understanding of the molecular mechanisms linking histone H3 core acetylation to transcriptional regulation.

      Strengths:

      (1) Site-specifically acetylated histone H3 proteins were chemically synthesized using a unique and rational peptide ligation strategy, representing a major technical strength of this study.

      (2) The in vitro transcription assays demonstrated that H3K56ac and H3K122ac increase the production of run-off transcripts, whereas H3K64ac has little effect on transcription efficiency. These findings highlight the distinct functional roles of individual acetylation sites within the histone H3 core domain.

      (3) The cryo-EM structures of nucleosomes containing either H3K56ac or H3K122ac revealed that H3 acetylation weakens histone-DNA interactions, providing a structural basis for the observed effects on transcription.

      Weaknesses:

      (1) Although the biochemical and structural data are convincing and sufficiently support the authors' conclusions, complementary cellular experiments would further strengthen the physiological relevance of the in vitro findings. While such experiments are not essential for supporting the main claims of the study, they would enhance the overall impact and biological significance of the work.

      (2) Although the authors demonstrate the structural consequences of individual H3 core acetylation events, the study does not investigate potential synergistic effects among multiple acetylated lysine residues within the H3 core domain. Consequently, the relationship between combinatorial acetylation patterns and their collective impact on RNA polymerase II-mediated transcription remains unclear.

    4. Reviewer #3 (Public review):

      This is a short and punchy manuscript that nicely summarises the 4 structures that are determined and provides a basis for the differences seen for acetylation sites shown for RNAPII activity.

      The authors build on previous biochemical work that determined the functional outcomes of H3 core acetylation, adapting an assay they have previously used extensively to investigate RNAPII transcription on nucleosomes and, indeed, even H3 N-terminal tail acetylation. This assay is as such well set up and has a wealth of confirmatory previous studies from this lab and the authors are careful not to overanalyse their results, leading to robust and well-considered results. The structures are determined to a high resolution, allowing the interpretation put forward about side chain orientations, with clear densities shown for the regions of interest.

      Further discussion or experiments would strengthen the conclusions further:

      (1) The conclusion on the role of H3K56Acetylation could be strengthened, especially as the results are somewhat counterintuitive. It is conceptually surprising that acetylation near the entry/exit DNA that destabilises this region also leads to a reduced stall propensity at the dyad but has a limited effect at SHL5? While it can be explained by the clash at the dyad pause being reduced, the more direct effect of DNA breathing amplification would be expected to have a larger effect at SHL 5. Indeed, the density for DNA at SHL5 appears to be weaker in Figure 2A, suggesting the entry/exit DNA flexibility is amplified past this region.

      Perhaps another assay that looks more directly at the flexibility of the entry/exit DNA would be useful, either through restriction enzyme-mediated cleavage or FRET (DNA ends and H2AK119 labels), providing stronger evidence of this effect. MNase is rather indirect and similar to the RNAPII assay itself.

      Similarly, were the authors surprised by the modest effect (less than 2-fold) in transcriptional pause at SHL 0 for the K122Ac? Presumably, based on the model in Figure 4, this would be expected to be the area with the largest effect? The results of K56Ac and K122Ac almost seem swapped to what would be expected in Figure 1H. Further discussion of this observation would be useful.

      (2) Could the local weakening of DNA, especially at the dyad, be observed in the cryo-EM structures? Perhaps comparison of local resolution estimation differences in this region compared to unmodified would be useful.

      (3) Caution should be taken, and discussion should include that the structural data presented is after extensive processing. Many nucleosome averaging classes were discarded in the 3D classification steps (nicely summarised in Table 1 as "particles for 3d classification" and "particles in final map"). Indeed, it is likely that higher DNA flexibility particles would be thrown away during this processing step. This can be observed for K56Ac DNA ordering, for example, in Supplementary Figure S4, yellow and cyan classes from the round of 3D classification look to be high resolution and have a higher order of DNA, so there has been some selection here. How was this done? While this is not fully quantifiable, it gives an idea of the extent of wrapping. We would suggest discussing the methodological limitations and showing the models after the first auto refinement to see if the features discussed on end flexibility and dan ordering are retained.

      (4) Di Cerbo et al. (reference 13) showed acetylation at K64 alters salt stability and affects transcription. Why do the authors think there is a discrepancy, albeit with different assays? Direct reference and discussion of this in the text should be included.

      (5) Why was H3.2 used, while this is relatively abundant in mouse cells, human protein was used, and this appears to be less common than H3.1 and H3.3. We are sure that the effect is not likely to be substantive on structure (as shown by the Kurumizaka lab previously), but should be addressed in the text

    1. eLife Assessment

      This important study provides a mechanistic view of how antibody affinity maturation can reshape encounter-state landscapes and association pathways, with implications for understanding HIV antibody maturation and vaccine design. The results are solid, supported by a coherent integration of adaptive molecular dynamics, Markov state modeling, SPR kinetics, mutagenesis, and double-mutant cycle analysis, although aspects of the kinetic validation, MSM-state robustness, and causal interpretation would benefit from further support. The work will be of interest to immunologists, structural biologists, and computational biophysicists.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript uses simulations and MSMs paired with experimental binding assays to examine the binding mechanisms of different antibodies to their targets. The authors argue that contacts in encounter complexes play an important role in determining the association rates and binding affinities that distinguish more mature antibodies from less efficacious antibodies from earlier in the maturation process.

      Strengths:

      The idea is interesting, and the combination of computational models and experiments is a good direction.

      Weaknesses:

      The manuscript focuses heavily on kinetics, but it is not clear whether the simulations recapitulate the relative rates of binding of the two antibodies. The relationship between the simulated binding behavior and the experimentally observed kinetic differences is therefore not fully established.

      The comparison of committor probabilities or fluxes between the two antibodies may not be appropriate. These properties are related to the barrier height the system has to cross to move forward vs back to the starting state, under the simplifying assumption that the properties of other states aren't critical. Even in this simplified case, the same flux or committor probability could occur with very different barrier heights, e.g., rates or transition probabilities.

      Some claims are presented in a very qualitative way that people who aren't experts in MSMs may have difficulty tying to the results in Figure 1.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript addresses an important and underexplored question: how affinity maturation alters antibody encounter-state landscapes rather than simply improving bound-state affinity. The authors combine adaptive MD, Markov State Models (MSMs), transition path theory, mutagenesis, SPR kinetics, and double-mutant cycle analysis into a coherent story.

      Strengths:

      This manuscript presents a compelling computational and experimental analysis of antibody affinity maturation in the HIV-1 DH270 lineage. The main finding is that somatic mutations reshape encounter-state pathways through glycan-mediated steering rather than simply stabilizing the final bound state. This is novel and potentially important for vaccine design. The combination of adaptive MD, MSMs, SPR kinetics, and double-mutant cycle analysis is a major strength.

      Weaknesses:

      The proposed sequence that somatic mutations cause glycan capture, which causes reorientation, which causes enhanced association, is based on correlation rather than direct causality.

      The four MSM states are not convincingly explained, and the robustness of these states is unclear.

      The productive collision surface area analysis needs more quantitative data.

      The coupling energy values are near the uncertainty range. Some conclusions about long-range communication networks appear stronger than the data justify. The data support coupling, but they do not necessarily support detailed mechanistic networks.

      The study investigates one lineage, one epitope class, and one viral system. Hence, the generalization is limited.

    4. Reviewer #3 (Public review):

      Summary:

      In this work, the authors set out to characterise how encounter states between antibodies and antigens evolve during affinity maturation through molecular dynamics simulations and Markov state modeling. They demonstrate how early glycan-mediated interactions increased association rates rather than modifying the final bound state.

      Strengths:

      The computational approach is backed up by experimental results and allows for visualising otherwise too short-lived association states, thus allowing to discriminate between different lineages.

      Weaknesses:

      The figures and captions are not always clear about what they are trying to show. The choice of CVs is not sufficiently discussed.

  2. Jul 2026
    1. eLife Assessment

      Muetter et al. provide an important argument that luminescence is a reliable, high-throughput alternative to colony-forming units (CFU) for super-MIC investigations, particularly when the quantity of interest is biomass. By examining 20 antimicrobials spanning 11 classes, the work shows that discrepancies between CFU and luminescence are often biological (filamentation, Viable But Not Culturable). The work provides a convincing view of how these three common measurements (luminescence, optical density, and CFU) relate to one another across a range of drug treatments, although testing on clinical isolates could be of further benefit.

    2. Reviewer #2 (Public review):

      Summary:

      In antibiotic research, accurately measuring decreases in bacterial populations is essential. The authors conducted a comprehensive evaluation of the luminescence assay, a commonly used but previously under-quantified method, benchmarking it against the gold-standard CFU counting approach. They found that luminescence measurements generally aligned with CFU results but sometimes reported slower decline rates for certain antimicrobials. These discrepancies were linked to differences in how the two methods capture biomass and colony formation, which vary with the antimicrobial's mechanism of action. The study demonstrates that luminescence assays can serve as a high-throughput alternative to labor-intensive CFU counting, provided their limitations are understood and corrected.

      Strengths:

      The authors developed a mathematical model to partially correct luminescence-based measurements, making the approach broadly applicable to several commonly used antibiotics. They also analyzed antibiotic-treated single-cell morphologies and linked filamentation to bulk luminescence signals. This analysis helped define the range of drug conditions under which luminescence assays provide reliable estimates of bacterial dynamics.

      They extensively evaluated the method using 20 antibiotics and one antimicrobial peptide, encompassing many of the most commonly used agents and experimental factors (e.g. treatment time) typically considered in antibiotic research.

      Comments on revised version:

      No further comments. The authors have adequately addressed my concerns.

    3. Reviewer #3 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability / carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorization of drug-specific assay behaviors.

      The study critically exposes flaws in the "gold standard" CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimize plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      In summary:<br /> Muetter et al. provide a compelling argument that luminescence is a reliable, high-throughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      Comments on revised version:

      The revised version addressed my comments well.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines how luminescence can be used to measure bacterial population dynamics during antimicrobial treatment by comparing it directly with optical density and colony counts. The authors aim to determine when luminescence reflects changes in population size and when it instead captures metabolic or physiological states induced by drug exposure. By generating parallel datasets under controlled conditions, the work provides a detailed view of how these three common measurements relate to one another across a range of drug treatments.

      Strengths

      The study is technically strong and thoughtfully designed. Measuring luminescence, optical density, and colony counts from the same cultures allows the authors to make clear and informative comparisons between methods. The data are compelling, and the analyses highlight both agreements and divergences in a way that is easy to interpret. The manuscript also succeeds in showing why these divergences arise. For example, the observation that filamentation and metabolic shifts can sustain luminescence even when colony counts drop provides valuable information on how different readouts capture distinct aspects of bacterial physiology. The writing is clear, the figures are effective, and the work will be useful for researchers who need high-throughput approaches to quantify microbial population dynamics experimentally.

      Weaknesses:

      The study also exposes some inherent limitations of luminescence-based measurements. Because luminescence depends on metabolic activity, it can remain high when cells are damaged or unable to resume growth, and it can fall quickly when drugs disrupt energy production, even if cells remain physically intact. These properties complicate interpretation in conditions that induce strong stress re-sponses or heterogeneous survival states.

      In addition, the use of drug-free plates for colony counts may overestimate survival when filamented or stressed cells recover once the antibiotic is removed, making differences between luminescence and colony counts harder to attribute to killing alone. Finally, while the authors discuss luminescence in the context of clinically relevant concentration ranges, the current implementation relies on engineered laboratory strains and does not directly demonstrate applicability to clinical isolates. These limitations do not detract from the technical value of the work but should be kept in mind by readers who wish to apply the method more broadly.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Luminescence limitations. We agree that the lack of a direct link between light intensity and a population property such as biomass or cell number is the main limitation of the luminescence method. To further emphasise this, we have expanded the Discussion in the revised manuscript.

      Drug-free plates. The use of drug-free plates is intentional. As we measure a time series, the question at each point is how many cells are alive at each time point. Cells that are stressed but viable at time t contribute correctly to the count at t. How long they survive under the respective treatment is captured by the subsequent timepoints.

      Filaments. Recovery of plated filamented cells should not inflate this estimate. A single plated filamentous cell is expected to yield either zero (death before division) or one single colony, regardless of in how many parts it separates, as all descendants are part of the same cluster. However, if the cells divide before plating, CFU can overestimate survival. Having that said, we have no indication that this occurred in our experiments, since in all observed discrepancies, CFU-based estimates were equal to or lower than those obtained from luminescence and the time cells spent in dilution was kept short.

      Clinical applicability. We agree with the reviewer that the method is not practical for ad-hoc pharmacodynamic studies of clinical isolates. What we instead provide is an E. coli-based model system to explore clinically relevant treatment conditions, which we address in the revised manuscript. We believe that constructing analogous bioluminescent model strains in other clinically relevant species would be a valuable direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability/carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorisation of drug-specific assay behaviours.

      The study critically exposes flaws in the “gold standard” CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimise plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      Weaknesses:

      The study is conducted exclusively using Escherichia coli. While E. coli is a standard model organism, the paper claims to evaluate luminescence as a generalisable high-throughput tool. Many of the discrepancies observed are driven by filamentation. However, distinct morphological responses occur in other critical pathogens (e.g., Staphylococcus aureus does not filament in the same way).

      The authors propose that luminescence data can be corrected using microscopyderived volume data to better align with CFU counts. The primary appeal of luminescence is high-throughput efficiency. If a researcher must perform timelapse microscopy to calculate cell volume changes to “correct” their luminescence data, the high-throughput advantage is lost.

      The paper argues that for ciprofloxacin, CFU underestimates viability because cells remain intact and impermeable to propidium iodide. While the cells are metabolically active and membrane-intact, if they cannot divide to form a colony (even after drug removal/dilution), their clinical relevance as “living” pathogens is debatable.

      Some other comments:

      The use of a population dynamical model to simulate filamentation effects is excellent. The finding that light intensity tracks volume ($\psi_V$) better than cell number ($\psi_B$) is a key theoretical contribution.

      The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      The use of bootstrapping to estimate rate distributions is appropriate and robust.

      Conclusion:

      Muetter et al. provide a compelling argument that luminescence is a reliable, highthroughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Generalisability. We agree that the alignments and divergences reported for specific drugs may not transfer directly to other species, which may elongate differently (e.g. cocci) or show different physiological responses to treatment. Constructing analogous model strains — for example based on S. aureus to cover a broader range of morphologies and clinically relevant species would therefore be an interesting follow-up project, and we have adjusted the Discussion to make this clearer. We nevertheless believe that the broader conclusions (larger cells emit more light) of the paper likely hold across species.

      Volume correction. We agree that requiring microscopy would undermine the high-throughput advantage of the luminescence assay. It was not our intention to propose this as a practical approach, nor to imply that the luminescence signal needs a correction. Taken on its own, the signal can be interpreted as the cumulative metabolic output of the population, which is closely linked to biomass, and that measure is valuable in itself for many applications. We used the volume correction only to demonstrate that luminescence tracks biomass more closely than cell number: by adjusting the luminescence distribution with the measured volume change, it moves towards the CFU distribution. We have revised the Discussion to prevent this from being misunderstood as a required step.

      Culturability vs. clinical relevance. We agree that the dynamics of culturable cells are highly relevant, especially in a clinical context. Our aim was to explain the observed differences between CFU and luminescence by highlighting that culturability and viability are not always identical, without implying that one measure is inherently superior to the other — we leave it to the reader to decide which metric best suits their needs.

      Linear elongation. The model assumes linear elongation for mathematical convenience, which, depending on the specific strain and drug mechanism, could be incorrect. Its purpose is to demonstrate that a shift of the mean cell volume to a new, higher equilibrium under treatment can cause an initial peak in the luminescence signal despite a declining population. This remains true for non-linear elongation models, though the shape, height and position of the peak may change. We have adjusted the Results to make this clearer.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors present luminescence as a practical measurement of population decline under antibiotic exposure. One aspect that could be clarified is how the method behaves when tolerance arises from phenotypic heterogeneity, such as the presence of small, metabolically quiet survivors. Because luminescence reflects metabolic activity and biomass, the signal will be dominated by metabolically active cells, making rare tolerant subpopulations difficult to detect. A short discussion of how luminescence performs in these heterogeneous scenarios, and whether complementary assays are needed to capture long-lived tolerant cells, would strengthen the manuscript.

      Yes, that is a valid concern and we thank the reviewer for raising this point.

      Heterogeneity in cell-specific luminosity alone does not bias population-level rate estimates. A bias can arise, however, when specific luminosity correlates with a second factor — most importantly, the decline rate under treatment.

      We agree with the reviewer’s suggestion that brighter cells plausibly die faster than tolerant, metabolically quiet ones. When one subpopulation dominates the light signal, we expect minimal bias, as the rate estimate primarily reflects that subpopulation. However, in a transition phase when both subpopulations contribute roughly equally to the light signal, luminescence likely overestimates the decline.

      We added a corresponding caveat to the Discussion (lines 581–583).

      (2) The manuscript shows that filamentation can influence ψ_I by altering biomass and metabolic activity independently of cell number. However, antibiotic exposure can also trigger other stress responses and metabolic shifts that change energy fluxes, redox balance, and biosynthetic activity. Since luminescence depends on metabolic state and substrate availability, these additional physiological transitions may also affect ψ_I in ways not directly tied to birth or death processes. It would be useful to comment on whether such responses, beyond filamentation, are likely to influence luminescence dynamics across different drug classes or treatment conditions.

      We thank the reviewer for raising this point and agree that there is no biological law strictly linking luminosity to a single population property such as biomass or cell number, and changes in the metabolism most likely affect Ψ<sub>I</sub> as well.

      Transitioning to a new metabolic steady state biases Ψ<sub>I</sub>; once the new steady state is reached, however, the rate estimate should no longer be affected.

      Looking across drug classes, drugs that primarily lyse cells (polymyxins and, to a lesser degree, beta-lactams targeting PBP1) did not show noticeable deviations between Ψ<sub>I</sub> and Ψ<sub>CFU</sub>, and — perhaps counterintuitively — neither did ribosome-inhibiting drugs.

      For the remaining cases, we were able to attribute part of the discrepancy between Ψ<sub>I</sub> and Ψ<sub>CFU</sub> to changes in biomass or loss of culturability, though drug-induced metabolic changes may also contribute to the residual differences.

      We clarify this in the Discussion (lines 569–579).

      (3) The authors quantify survival using colony counts on drug-free medium. Because filamentation can be a reversible state that persists during antibiotic exposure, plating on drug-free medium may capture recovery potential rather than in-treatment viability. Filamented or stressed cells that cannot divide in the presence of a drug may nevertheless form colonies once the drug is removed. Clarifying how this recovery step affects ψ_CFU would help readers interpret differences between luminescence-based and colony-based measurements, especially in cases where transient tolerant states are present.

      We thank the reviewer for raising this point.

      Our CFU assay estimates the number of culturable cells at each time point; the rate Ψ<sub>CFU</sub> is then inferred from how this number changes across time points. Plating on drug-free medium is intentional, as it maximises the probability that a culturable cell is detected at each snapshot. Whether those cells would have continued dividing or died under continued treatment is captured by the subsequent time points.

      Filamentation interacts with the probability of colony formation in several, partly opposing ways:

      (1) It can increase the death rate, as for ceftazidime and cefepime, which is part of the kill effect captured by Ψ<sub>CFU</sub>;

      (2) Entanglement between filaments may reduce the number of colonies per plated bacterium;

      (3) Conversely, if a filament divides upon drug removal, its fragments form a cluster that — stochastically — is very likely to produce one (but not multiple) colony.

      The only scenario in which CFU could overestimate bacterial density is if a filament separates into individual cells in the liquid phase before plating; we have no indication that this occurred in our experiments.

      We addressed this concern in our response to the public comment.

      (4) A brief discussion comparing luminescence to fluorescent reporter systems could be helpful. Fluorescent proteins typically require a chromophore maturation step before becoming detectable, which introduces a delay between the underlying cellular event and the appearance of the signal. In contrast, as far as I understand, lux reporters emit light immediately once the enzymatic components and substrates are present, without a maturation stage. Highlighting this distinction may help readers understand why luminescence is well-suited for tracking rapid changes in population physiology under antibiotic exposure. However, the manuscript also notes that luminescence can lag slightly behind very rapid killing (particularly for AMPs), but the temporal dynamics of signal shutdown are not explored in detail. Because lux reflects metabolic activity rather than viability, a short delay between irreversible damage and the loss of light is biologically expected. It may help readers if the authors could expand on the mechanism underlying this delay in order to clarify when ψ_I is likely to track true biomass decline and when residual metabolic activity might mask early killing events.

      On fluorescent reporters:

      We thank the reviewer for this suggestion.

      Under some conditions, change rates can also be measured using fluorescence, provided the number of fluorescent molecules per bacterium remains constant. This requires a balance between production, maturation, degradation and dilution, which is only established if the growth rate and conditions remain constant over a sufficiently long period (typically hours).

      For measuring population decline, however, the key issue is that cell death does not inactivate fluorescent proteins: once matured, they emit independently of the cell’s metabolic state and decay only with the protein’s half-life, which is typically slower than the kill rates of interest.

      We added a clarification to the Introduction (lines 58–60).

      On the lux signal lag:

      We thank the reviewer for raising this point. The short lag between luminescence and CFU decline could in principle arise from two mechanisms: (i) luminescence declining more slowly than the actual cell number (residual light from dead cells), or (ii) CFU declining more steeply than the actual cell number (damaged but still viable cells failing to form colonies).

      Mechanism (i) splits into two sub-cases:

      (i.a) Dead but impermeable — the lux reaction could in principle continue for a short while if enough components are retained in the cell. However, a metabolically active, impermeable cell is difficult to classify as dead in the first place, making this scenario conceptually awkward.

      (i.b) Dead and permeable (lysed) — the lux components dilute into the medium, and by mass-action the reaction rate should drop rapidly (though not instantly). Any residual signal after lysis should therefore be short-lived.

      Mechanism (ii) — damaged (e.g. permeable) cells may be particularly sensitive to plating on agar (e.g. due to oxidative stress), resulting in a declining probability of colony formation.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse, making (i.b) and/or (ii) the likely explanations. Based on our experimental data, we cannot distinguish between these possibilities and therefore limit ourselves to reporting the observed discrepancy.

      We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded the discussion there.

      (5) In lines 85–89, the authors state that “high-throughput OD and luminescence measurements at sub-MIC concentrations provide valuable insights into drug effects on growth rates, [but] the super-MIC range is clinically more relevant,” and they present luminescence as a way to investigate super-MIC population dynamics. While super-MIC behaviour is indeed important for pharmacodynamics and resistance evolution, it is not clear that the specific luminescence implementation used here has direct clinical relevance. The study relies on a chromosomally integrated reporter in a laboratory strain, and the manuscript does not demonstrate that this approach can be applied to clinical isolates or diagnostic workflows. It may be helpful to moderate the claim of “clinical relevance” and frame the method more clearly as a high-throughput experimental tool that can inform clinically relevant questions, rather than as an assay ready for clinical application.

      We agree and have moderated the framing accordingly (lines 100–104).

      Reviewer #2 (Recommendations for the authors):

      (1) The conclusions regarding “biomass vs. cell number” may not apply equally to non-rod-shaped bacteria or species with different stress responses. The authors must explicitly discuss this limitation in the Discussion.

      The broad conclusion that bigger cells emit more light likely holds across morphologies, since it rests on the principle that more cellular material means more metabolic activity and therefore more light. The quantitative relationship between cell size and luminosity, however, may differ across species, shapes and conditions, for two reasons. First, chromosome copy number: whether drug-induced morphological changes are accompanied by chromosome replication and therefore an increase in lux operon copy number — varies across species and drug mechanisms. Second, the surface-to-volume ratio likely modulates mass-specific metabolism; some morphological changes preserve it (e.g. purely lateral elongation) while others do not.

      The more specific conclusions about which drug classes produce alignment or divergence between CFU and luminescence may also not transfer directly, as drug mechanisms can act differently across species.

      We already note this limitation in the Discussion (lines 594– 597) and have expanded the wording there.

      (2) The manuscript should clarify that luminescence is a superior metric for biomass without correction, rather than framing the volume correction as a necessary step to mimic CFU. The divergence should be embraced as a feature (biomass tracking), not a bug that needs fixing via labor-intensive microscopy.

      We agree with the framing and will make it clearer; it was actually our intention to clarify which method does what, rather than judge one as better or worse.

      We removed the “correction” sentence from the Discussion to make this clearer.

      (3) The authors should refrain from definitively stating CFU “underestimates” viability and instead use more precise terminology, such as “reproductive capability” vs. “metabolic integrity.”

      We agree with the reviewer that measuring culturability is a property, not a flaw, of CFU. Our intention was to emphasise that when CFU is used as a proxy for viability (which it often is), it can yield lower values than the actual number of survivors. We tried to make that distinction explicit in the manuscript (e.g. in lines 317–322).

      We would also like to note that in the case of antimicrobial carryover, CFU can genuinely underestimate culturability itself, not only viability.

      Regarding the suggested reproductive capability vs. metabolic integrity framing: we agree that metabolism and luminescence are closely linked. What held us back from drawing that link directly is that metabolism is hard to quantify, being the cumulative output of a diverse set of processes.

      (4) The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      Linear elongation is a mathematically convenient simplification whose only purpose in the model is to allow the population to converge to a new equilibrium volume under treatment. Assuming constant volume-specific luminosity, we showed that this produces an initial peak in light intensity before the signal declines in parallel with Ψ<sub>B</sub>. The exact shape, height and position of this peak depend on the volume growth model used, but the qualitative pattern — peak followed by parallel decline — holds for other growth models as well. We now clarify this in lines 230–235.

      (5) The authors suggest the carryover effect is due to a delay between cell death and cessation of luminescence. This “lag time” is a critical physical constraint of the lux system (likely related to ATP depletion or enzyme decay) and should be quantified or discussed in more detail as a fundamental “speed limit” for the assay.

      The origin of the lag between luminescence and CFU is an interesting question, but one we cannot definitively answer. We can, however, discuss the potential mechanisms:

      A dead but impermeable cell could in principle continue to emit residual light for some time. We note, though, that calling a metabolically active, impermeable cell “dead” is a question of definition we would rather not discuss here.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse. Under lysis, the lux components dilute quickly into the medium, and by mass-action the reaction rate should drop rapidly — though not necessarily instantaneously.

      A plausible alternative to a delayed cessation of the light signal is that the probability of colony formation drops rapidly after permeabilisation, for example because permeable cells are sensitive to oxidative stress when plated on agar.

      Based on our experimental data we cannot distinguish between these mechanisms, so we limit ourselves to reporting the observed discrepancy. We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded on the candidate mechanisms there.

      Additional revisions

      Beyond the changes prompted by the reviewers’ comments, we made the following revisions to the supplementary information:

      We corrected the Λ matrix (converted row 2, col 4 from 0 → 2)

      We removed the line numbering

    1. eLife Assessment

      This valuable study introduces MULTI i<sup>2</sup>, a robust and high-throughput method to measure Plasmodium falciparum viability in the presence of drugs. This new assay offers significant time savings over the traditional Parasite Reduction Rate (PRR) assay and should enable faster screening of drug combinations, which is urgently needed in the field. The assay is well validated, with convincing data showing it can reproduce known drug interactions and identify new interaction patterns.

    2. Reviewer #1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRRv2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRRv2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRRv2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      There are a number of areas for improvement:

      (1) Many antimalarials have quite specific times of action. Are these MULTI-i2 assays, and the comparator PRRv2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      (2) The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULTI-i2 method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      (3) It would be helpful for authors to provide some indication of the cost comparison between the PPRv2 and MULTI-i2.

      (4) Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      (5) The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

    3. Reviewer #2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRRv2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i2 assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i2 assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i2 assay.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i2 assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRRv2 assay?

      The addition of an inducible element is an improvement of their earlier lacZ/β-galSENSOR (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRRv2, they fail to compare it to their own non-inducible lacZ/β-galSENSOR system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved? How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

    4. Reviewer #3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULTI-i2, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i2 assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULTI-i2 assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULTI-i2 provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULTI-i2 methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1) The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      (2) Related to that above, how would MULTI-i2 perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      (3) Given the stated cost and labor efficiency of MULTI-i2, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i2 method more attractive. In particular, it would be nice to see if one could use MULTI-i2 for studies of triple combinations as enthusiastically suggested.

      (4) Throughout the manuscript, the authors claim that MULTI-i2 is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

    5. Author response:

      Reviewer 1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we will also explicitly state this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we will mention and describe alternative approaches for assessing parasite viability. This will also include the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We will revise the text to explicitly mention the use of dual straining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al, 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer 2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We will clarify this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We will add this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We will revise the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      Reviewer 3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1): The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We will add more explanations to the Discussion including the strengths and limitations.

      (2) Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      (3) Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      (4) Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. eLife Assessment

      This important study combines anatomical tracing, tissue clearing, and functional manipulations to demonstrate lateralized brainstem control of hepatic glucose metabolism and identify a site of sympathetic nerve crossover supplying the liver. The evidence supporting the anatomical organization of hepatic sympathetic innervation is compelling, and the functional studies provide solid support for a role of asymmetric sympathetic outflow in regulating glucose homeostasis. While some uncertainty remains regarding the contribution of sensory innervation and the extent to which these findings generalize beyond mice, the work provides an invaluable advance in understanding neural regulation of liver metabolism.

    2. Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether hepatic glucose metabolism is differentially regulated by the left and right sides of the LPGi and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi, which were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, and in changes in protein expression in the liver lobes. These data suggested lobe-specific modulation of HGP. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) Image clarity was improved in some cases, but not in others. For example, Figure 3I, showing c-Fos expression, is not convincing due to the image quality and lack of orientation.

      (2) The methods section states that 8-weeks-old male mice were used in the experiments without specifying the experiments (e.g., brain injection with AAVs or PRV organ inoculation). The authors should include these details.

      (3) The authors should use the exact location of pre- and postganglionic neurons as they often refer to neurons in the sympathetic chain. Their findings should be compared with the existing literature on the location of preganglionic cells.

      (4) Figure legends should be revised and matched with the text.

    3. Reviewer #4 (Public review):

      Summary of General Strengths & Weaknesses:

      The studies here are highly informative for anatomical tracing and sympathetic nerve function in the liver in relation to glucose levels, but because they are conducted in a single species, it is challenging to translate them to humans or determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies provides mechanistically informative. Denervation studies lack proper controls, and sensory innervation in the liver is overlooked.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also labels a portion of sensory fibers that need to be ruled out in whole-mount imaging data.

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases in tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though that is clearly assumed. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      Comments on the revised version.

      Across all reviewer comments, the revised resubmission has adequately addressed all concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Wang et al. reports the potential involvement of an asymmetric neurocircuit in the sympathetic control of liver glucose metabolism.

      Strengths:

      The concept that the contralateral brain-liver neurocircuit preferentially regulates each liver lobe may be interesting.

      Weaknesses:

      However, the experimental evidence presented did not support the study's central conclusion.

      We thank the reviewer for recognizing the conceptual novelty of our work and for constructive comments aimed at enhancing its rigor and clarity. In response, we carried out targeted experiments to address the points raised, including: (i) further characterization of LPGi projections to vagal and sympathetic circuits; (ii) evaluation of potential pancreatic involvement; and (iii) validation of the specificity of chemogenetic activation within the proposed circuit. All new experiments, figures, and text have been incorporated, and corresponding revisions are highlighted for ease of review.

      (1) Pseudorabies virus (PRV) tracing experiment:

      The liver not only possesses sympathetic innervations but also vagal sensory innervations. The experimental setup failed to distinguish whether the PRV-labeling of LPGi (Lateral Paragigantocellular Nucleus) is derived from sympathetic or vagal sensory inputs to the liver.

      Thank you for raising this important point. We fully agree that the liver receives both sympathetic and vagal sensory innervation, and we acknowledge that PRV-based tracing alone does not definitively distinguish between these two pathways. This represented a limitation of the original experimental design.

      Based on established anatomical literature as well as our experimental observations, vagal sensory neuron cell bodies reside in the nodose ganglion (NG), and their central projections terminate predominantly in the nucleus of the solitary tract (NTS) (Nature. 2023;623(7986):387-396; Curr Biol. 2020;30(20):3986-3998.e5.), which is located in the dorsomedial medulla. In contrast, the LPGi, together with other sympathetic-related nuclei, is predominantly distributed in the ventral medulla (Cell Metab. 2025;37(11):2264-2279.e10; Nat Commun. 2022;13(1):5079).

      To determine whether the LPGi contains neurons that modulate the liver via vagal sensory pathways, we performed two complementary experiments.

      First, we conducted CGRP immunohistochemistry on brainstem sections, using the NTS, a well-established visceral sensory centre, as a positive control. While abundant CGRP-positive cell bodies were detected in the NTS as expected, few to no CGRP-positive cell bodies were observed in the LPGi (Figure S1G). These results strongly support that the LPGi neurons labeled in our PRV tracing predominantly belong to sympathetic efferent circuits rather than vagal sensory pathways.

      Second, to examine whether LPGi neurons send axonal projections to sensory ganglia, we injected hSyn-Cre combined with DIO-Axon-EGFP into the LPGi and examined both the dorsal root ganglia (DRG) and nodose ganglia (NG). No Axon-EGFP-positive signals were detected in either ganglion (Figures S1H-S1J), indicating that LPGi neurons do not directly innervate sensory ganglia. In other words, PRV cannot retrogradely trace to the LPGi via the NG or DRG.

      These additions have been incorporated into the revised manuscript, with the Result 1 clearly documenting that these findings confirm that the LPGi specifically regulates sympathetic, rather than vagal sensory, inputs to the liver.

      (2) Impact on pancreas:

      The celiac ganglia not only provide sympathetic innervations to the liver but also to the pancreas, the central endocrine organ for glucose metabolism. The chemogenetic manipulation of LPGi failed to consider a direct impact on the secretion of insulin and glucagon from the pancreas.

      Thank you for this important comment. We agree that the celiac ganglia (CG) provide sympathetic innervation not only to the liver but also to the pancreas, which plays a central role in glucose homeostasis through the secretion of both insulin and glucagon. Therefore, the potential pancreatic implications associated with LPGi chemogenetic manipulation are worth careful consideration.

      To address this concern, we measured circulating glucagon and insulin levels following chemogenetic manipulation of the LPGi<sup>GAD1</sup> neurons. We found that neither glucagon nor insulin levels changed significantly under our experimental conditions, which indicated that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated by changes in pancreatic hormone secretion (Figure S2G).

      These additions have been incorporated into the revised manuscript, with the Result 2 clearly documenting that these findings confirm that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated indirectly via altered pancreatic endocrine output.

      (3) Neuroanatomy of the brain-liver neurocircuit:

      The current study and its conclusion are based on a speculative brain-liver sympathetic circuit without the necessary anatomical information downstream of LPGi.

      Thank you for raising this important point. A clear anatomical definition of the downstream pathways linking the brain to the liver was essential for interpreting the proposed brain-liver sympathetic circuit.

      The present study (Figure 4A) provides direct anatomical evidence supporting the organization of the brain–liver sympathetic neurocircuit. These observations are consistent with our recent detailed characterization of the brain-liver sympathetic circuit published in Cell Metabolism (Cell Metab. 2025;37(11):2264–2279). In that study, we showed that LPGi GABAergic neurons inhibit GABAergic neurons in the caudal ventrolateral medulla (CVLM). Disinhibition of CVLM reduced GABAergic suppression of rostral ventrolateral medulla (RVLM) neurons, which are key excitatory drivers of sympathetic tone. RVLM neurons project to sympathetic preganglionic neurons in the sympathetic chain (Syc). These neurons synapse with postganglionic sympathetic neurons in ganglia such as the celiac-superior mesenteric ganglion (CG-SMG). Postganglionic sympathetic fibers then innervate the liver, releasing norepinephrine (NE) to activate hepatic β<sub>2</sub>-adrenergic receptors and stimulate hepatic glucose production (HGP).

      Together, these data establish a coherent anatomical basis for the proposed brain-liver sympathetic pathway and clarify the downstream organization relevant to the functional experiments presented in figure 4A and Author response image 1..

      Author response image 1.

      Tracing scheme (Left) and whole-mount imaging (Right) of PRV-labeled brain-liver neurocircuit. Scale bars, 3,000 (whole mount) or 1,000 (optical sections) μm.

      (4) Local manipulation of the celiac ganglia:

      The left and right ganglia of mice are not separate from each other but rather anatomically connected. The claim that the local injection of AAV in the left or right ganglion without affecting the other side is against this basic anatomical feature.

      Thank you for raising this important anatomical point. We fully acknowledge that the left and right CG in mice are interconnected, and that unilateral viral injection could theoretically affect the contralateral side. The CG-SMG complex serves as a major sympathetic hub that regulates visceral organ functions. Recent transcriptomic, anatomical, and functional studies have revealed that the CG-SMG is not a homogeneous structure but is composed of molecularly and functionally distinct neuronal populations. These populations exhibit specialized projection patterns and regulate different aspects of gastrointestinal physiology, supporting a model of modular sympathetic control. (Nature. 2025 Jan;637(8047):895-902). Therefore, we were aware of this phenomenon during the initial stages of these experiments.

      To minimize unintended spread to the contralateral CG, we took two complementary approaches.

      First, we optimized the injection strategy by using an extremely small injection volume (100 nL per site), with a very slow infusion rate (50 nL/min), and fine glass micropipettes. With these refinements, contralateral viral spread was rarely observed.

      Second, and importantly, all animals included in the final analyses were subjected to post hoc anatomical verification. After completion of the experiments, CGs were collected, sectioned, and examined for viral expression. As shown in Supplementary Figure 5F, only mice in which viral expression was strictly confined to the targeted CG, with no detectable infection in the contralateral ganglion, were included in the presented data.

      Together, these measures ensure that our local manipulation of the intended CG produced the reported effects. We have revised the Methods section to more explicitly detail these technical precautions, and the legend for Figure S5F clearly states its role in validating injection specificity.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether the left and right LPGi differentially regulate hepatic glucose metabolism and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, as well as changes in protein expression in the liver lobes. These data suggested modulation of HGP (hepatic glucose production) in a lobe-specific manner. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      We thank the reviewer for the thorough and constructive evaluation of our manuscript. In direct response, we undertook comprehensive revisions to enhance the rigor and clarity of the study, including: (i) correcting ambiguous or misleading terminology about anatomical resolution and sympathetic circuit organization; (ii) expanding the Methods section with complete experimental details, improved image presentation, and explicit justification of our viral and genetic approaches; and (iii) strengthening data interpretation by addressing issues related to sparse PRV labeling, projection heterogeneity, and the functional implications of double-labeled neurons.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) The wording/terminology used in the manuscript is misleading, and it is not used in the proper context. For instance, the goal of the study is "to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism..." (see abstract); however, the authors focus on the brainstem (a single structure without hemispheres). Similarly, symmetric is not the best word for the projections.

      We thank the reviewer for raising these critical points regarding terminology and conceptual framing. We acknowledge that certain phrases in our original manuscript may have been overly broad or ambiguous, particularly in describing the scope of sympathetic heterogeneity and the specificity of neural projections. Due to practical constraints and the scope of our study, our investigation focused on the brainstem, which represents the final common pathway for these lateralized commands. We acknowledge that terms referring to the cerebral hemispheres do not accurately describe our study. We have revised the manuscript to ensure accurate and consistent terminology.

      Below are specific examples:

      Original 1: This study aims to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism and localize the site of sympathetic crossover to the liver.

      Revised 1: “This study aimed to determine whether the central nervous system exerts lateralized control over hepatic glucose metabolism and to localize the site of peripheral sympathetic crossover to the liver.”

      Original 2: These findings demonstrate that the brain exerts lobe-specific, lateralized control of hepatic glucose metabolism via symmetric brain-liver sympathetic pathways.

      Revised 2: “These findings demonstrate that the brainstem can exert lobe-specific, lateralized control of hepatic glucose metabolism via bilaterally projecting brain–liver sympathetic pathways.”

      Original 3: The cerebral hemispheres exhibit pronounced functional asymmetry, [1,2] a phenomenon traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation.[3,4]

      Revised 3: “Pronounced functional lateralization within the central nervous system (CNS) is a well-documented phenomenon, [1,2] traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation [3,4].

      (2) Sparse labeling of liver-related neurons was shown in the LPGi (Figure 1). It would be ideal to have lower magnification images to show the area. Higher quality images would be necessary, as it is difficult to identify brainstem areas. The low number of labeled neurons in the LPGi after five days of inoculation is surprising. Previous findings showed extensive labeling in the ventral brainstem at four days post-inoculation (Desmoulins et al., 2025). Unfortunately, it is not possible to compare the injection paradigm/methods because the PRV inoculation is missing from the methods section. If the PRV is different from the previously published viral tracers, time-dependent studies to determine the order of neurons and the time course of infection would be necessary.

      We sincerely thank the reviewer for these detailed and constructive comments regarding the PRV tracing experiments. We fully agree that careful presentation and interpretation of the anatomical data are essential for ensuring rigor and transparency. We address each point in detail below.

      (1) Image magnification and anatomical context of LPGi labeling

      We agree that the original images did not sufficiently convey the broader anatomical context of the LPGi. Due to fluorescence quenching in previous sections, we repeated the PRV retrograde tracing experiment and performed statistical analysis. In the revised manuscript, we replaced the original panels in Figure 1 and Figure S1 with new images that include lower-magnification overviews of the brainstem, alongside higher-magnification views of the LPGi (Figure 1). These images clearly delineate the LPGi with respect to established anatomical landmarks and atlas boundaries. Image contrast and resolution were optimized to allow unambiguous identification of PRV-labeled neurons and surrounding structures.

      (2) Sparse LPGi labeling at 5 days post-injection and methodological details

      We apologize for the omission of the detailed PRV injection protocol in the original Methods section. We deliberately used small-volume, local injections (1 µL per liver lobe) to minimize viral spread and to restrict labeling to circuits specifically connected to the targeted hepatic region. This sparse labeling was consistent with the use of small, spatially restricted injections designed to minimize off-target spread and preferentially label higher-order upstream neurons. This information has now been added, including the PRV strain, viral titer, injection volume, precise injection coordinates, and surgical procedures. All new figures, legends, and Method details have been incorporated, with changes clearly highlighted for ease of review.

      These additions have been incorporated into the revised manuscript, in Figure 1B-D, Figure S1C-F. Furthermore, we also added details of the Methods.

      (3) Not all LPGi cells are liver-related. Was the entire LPGi population stimulated, or was it done in a cell-type-specific manner? What was the strain, sex, and age of the mice? What was the rationale for using the particular viral constructs?

      We thank the reviewer for this insightful and important question. We agree that not all neurons within the LPGi are liver-related, and we apologize that our rationale was not clearly articulated in the original manuscript.

      (1) Our decision to target GABAergic neurons in the LPGi using GAD1-Cre mice was based on prior experimental evidence rather than an assumption about the entire LPGi population. In our previous study (Cell Metab. 2025;37(11):2264-2279.e10), we performed single-cell RNA sequencing on retrogradely labeled LPGi neurons following liver tracing. These analyses revealed that the majority of liver-projecting LPGi neurons are GABAergic in nature. Based on these findings, we chose to selectively manipulate GABAergic neurons in the LPGi rather than the entire LPGi neuronal population, to achieve greater cellular specificity and to minimize potential confounding effects arising from heterogeneous neuron types within this region. We regret that this rationale was not clearly described in the original submission and have now revised the manuscript to explicitly state this reasoning (Results section 2, paragraph 2: “Prior single-nucleus RNA sequencing and immunofluorescence analyses demonstrated that liver-projecting LPGi neurons are predominantly GABAergic.”).

      (2) In addition, we apologize for the omission of mouse strain, sex, and age information in the Methods section. These details have been fully added.

      (3) We selected AAV-based viral vectors, specifically the AAV9 serotype, due to their well-established efficiency in transducing neurons in the brainstem, relatively low toxicity, and widespread use in circuit-level chemogenetic and optogenetic studies. When combined with Cre-dependent viral constructs in GAD1-Cre mice, this approach enabled selective and reliable manipulation of LPGi GABAergic neurons.

      (4) The authors should consider the effect of stimulation of double-labeled neurons (innervating more than one lobe) and potential confounding effects regarding other physiological functions.

      We thank the reviewer for raising this important point. We agree that neurons innervating more than one liver lobe could, in principle, introduce potential confounding effects and may reflect higher-order integrative autonomic neurons.

      This consideration is consistent with a key finding of the cited study: the CG-SMG contains molecularly distinct sympathetic neuron populations (e.g., RXFP1<sup>+</sup> vs. SHOX2<sup>+</sup>) that exhibit complementary organ projections and separate, non‑overlapping functions. Specifically, RXFP1<sup>+</sup> neurons innervate secretory organs (pancreas, bile duct) to regulate secretion, while SHOX2<sup>+</sup> neurons innervate the gastrointestinal tract to control motility. This functional segregation supports the concept of specialized autonomic modules rather than a uniform, “fight-or-flight” response, reinforcing the need for careful interpretation of circuit-specific manipulations. (Nature. 2025;637(8047):895-902; Neuron. 2026;114(3):463-478.e7).

      In our PRV tracing experiments, the proportion of double-labeled neurons was relatively small, suggesting that the majority of labeled LPGi neurons preferentially associate with individual hepatic lobes. Nevertheless, we recognize that activation of this minority population could contribute to broader physiological effects beyond strictly lobe-specific regulation. We have therefore added a paragraph in the second paragraph of the Discussion (Paragraph 2: “A small subset of LPGi neurons was double-labeled after bilateral PRV injections, suggesting a fraction of these neurons projects bilaterally to both sides of the liver. Such neurons may support interlobar coordination.”).

      (5) The authors state that "central projections directly descend along the sympathetic chain to the celiac-superior mesenteric ganglia". What they mean is unclear. Do the authors refer to pre-ganglionic neurons or premotor neurons? How does it fit with the previous literature?

      We thank the reviewer for pointing out this imprecise wording. We agree that the original phrasing was anatomically inaccurate and potentially confusing. The pathways we intended to describe involve brainstem premotor neurons that project to sympathetic preganglionic neurons in the spinal cord. These preganglionic neurons then innervate neurons in the CG-SMG, which in turn provide postganglionic input to the liver.

      We have revised the manuscript to clearly distinguish premotor from preganglionic neurons (Results section 4, paragraph 1: “Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the spinal cord send descending fibers through the sympathetic chain (SyC) to innervate postganglionic neurons in the CG-SMG (Figure 4A). Further whole-mount TH immunostaining showed that TH-positive sympathetic cell bodies within the CG-SMG project to the liver along the hepatic vasculature (Figure 4B and Figure S5F). These observations suggest that the nerve bundles likely decussate at the porta hepatis before entering the individual hepatic lobes.”).

      (6) How was the chemical denervation completed for the individual lobes?

      We thank the reviewer for raising this important methodological concern. We agree that potential diffusion of 6-OHDA is a critical issue when performing lobe-specific chemical denervation, and we apologize that our original description did not sufficiently clarify how this was controlled.

      In the revised Methods section, we provided a detailed description of the denervation procedure, including the injection volume and concentration of 6-OHDA, as well as the physical separation and isolation of individual hepatic lobes during application to minimize diffusion to adjacent tissue.

      To directly assess the specificity of the chemical denervation, we included immunofluorescence and Western blot analyses demonstrating a selective reduction of sympathetic markers in the targeted lobe (Figure 3C), with minimal effects on non-targeted lobes. These results support the effectiveness and relative spatial confinement of the 6-OHDA treatment under our experimental conditions.

      We thank the reviewer for highlighting this point, which has helped us improve both the clarity and rigor of the manuscript.

      (7) The Western Blot images look like they are from different blots, but there are no details provided regarding protein amount (loading) or housekeeping. What was the reason to switch beta-actin and alpha-tubulin? In Figures 3F -G, the GS expression is not a good representative image. Were chemiluminescence or fluorescence antibodies used? Were the membranes reused?

      We thank the reviewer for this careful and detailed evaluation of the Western blot data. We apologize that insufficient methodological detail was provided in the original submission.

      (1) We would like to clarify that the protein bands shown within each panel were derived from the same membrane. To improve transparency, we provided full, uncropped images of the corresponding membranes in the supplementary materials. In addition, detailed information regarding protein loading amounts, gel conditions, and housekeeping controls has also been added to the Methods section.

      (2) The use of different loading controls (β-actin or α-tubulin) reflects a technical consideration rather than an experimental inconsistency. In our experiments, the molecular weight of the TH (62kDa) was too close to that of α-tubulin (55kDa), and β-actin (42kDa) was therefore used to avoid band overlap and to ensure accurate quantification.

      (3) Regarding the GS signal shown in Figures 3F–G, we agree that the original representative image was suboptimal. This appears to be related to antibody performance rather than sample quality. To address this, we repeated the Western blot from Figures 3F–G using a newly validated antibody. The original tissue samples had been aliquoted and stored at −80 °C, allowing reliable re-analysis.

      (4) All Western blot experiments were detected using chemiluminescence, and membrane stripping and reprobing procedures are now explicitly described in the Methods section.

      We thank the reviewer for highlighting these issues, which significantly improve the rigor and clarity of our data presentation. All new figures and legends have been incorporated, with changes clearly highlighted for ease of review.

      (8) Key references using PRV for liver innervation studies are missing (Stanley et al, 2010 [PMID: 20351287]; Torres et al., 2021 [PMID: 34231420]; Desmoulins et al., 2025 [PMID: 39647176]).

      We thank the reviewer for pointing out these important and highly relevant references that were inadvertently omitted in our initial submission. The studies by Stanley et al. (Proc Natl Acad Sci U S A, 2010), Torres et al. (Am J Physiol Regul Integr Comp Physiol, 2021), and Desmoulins et al. (Auton Neurosci, 2025) represent key PRV-based retrograde tracing work that has mapped central neural circuits innervating the liver and thus provide essential context for our anatomical analyses.

      We agree that the inclusion of these studies is necessary to properly situate our findings within the existing literature. Accordingly, we incorporated citations to these references in the revised manuscript and discussed their relationship to our results.

      Reviewer #3 (Public review):

      Summary:

      This study found a lobe-specific, lateralized control of hepatic glucose metabolism by the brain and provides anatomical evidence for sympathetic crossover at the porta hepatis. The findings are particularly insightful to the researchers in the field of liver metabolism, regeneration, and tumors.

      Strengths:

      Increasing evidence suggests spatial heterogeneity of the liver across many aspects of metabolism and regenerative capacity. The current study has provided interesting findings: neuronal innervation of the liver also shows anatomical differences across lobes. The findings could be particularly useful for understanding liver pathophysiology and treatment, such as metabolic interventions or transplantation.

      Weaknesses:

      Inclusion of detailed method and Discussion:

      We sincerely thank the reviewer for the positive and constructive feedback, which significantly enhances both the methodological rigor and the broader biological interpretation of our study. In direct response, we revised the Discussion to elaborate on the potential physiological advantages of a lateralized and lobe-specific pattern of liver innervation. Furthermore, we expanded the Methods section to include a comprehensive description of the quantitative analysis applied to PRV-labeled neurons. Together, these revisions strengthened the manuscript’s clarity, depth, and relevance to researchers in hepatic metabolism, regeneration, and disease.

      (1) The quantitative results of PRV-labeled neurons are presented, and please include the specific quantitative methods.

      We thank the reviewer for this helpful suggestion. We have added a detailed description of the quantitative methods used to analyze PRV-labeled neurons in the revised Methods section. We have now provided detailed information in the Methods section, including the criteria used for cell counting, the anatomical boundaries of the brain regions analyzed, the delineation of regions of interest, and the normalization procedures applied to derive the reported neuron counts. These additions have been incorporated into the revised Methods, with all changes clearly indicated for ease of review.

      (2) The Discussion can be expanded to include potential biological advantages of this complex lateralized innervation pattern.

      We appreciated the reviewer’s suggestion. We have expanded the Discussion to include a paragraph addressing the potential biological significance of lateralized liver innervation. We highlight that this asymmetric organization could allow for more precise, lobe-specific regulation of hepatic metabolism, enable integration of distinct physiological signals, and potentially provide robustness against perturbations. The additional discussion content has been highlighted in the revised version as indicated (Discussion section, paragraph 3: “Bilateral LPGi activation produced additive effects, indicating that both sides of the brainstem can cooperatively regulate hepatic metabolism in a spatially segregated manner. This pattern suggests that hepatic glucose output can be modulated in a lobe-specific, rather than uniform whole-organ, manner.”).

      Reviewer #4 (Public review):

      Summary:

      The studies here are highly informative in terms of anatomical tracing and sympathetic nerve function in the liver related to glucose levels, but given that they are performed in a single species, it is challenging to translated them to humans, or to determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies is mechanistically informative. Denervation studies lack appropriate controls, and the role of sensory innervation in the liver is overlooked.

      We sincerely appreciate the reviewer's thoughtful evaluation and fully agree that findings derived from a single-species model must be interpreted with caution in relation to human physiology. In direct response, we revised the manuscript to explicitly clarify that all experimental data were obtained in mice and to provide a discussion of the limitations regarding direct extrapolation to humans. Concurrently, we expanded the Discussion section by integrating our findings with recent human and translational studies, including a multicenter clinical trial demonstrating that catheter-based endovascular denervation of the celiac and hepatic arteries significantly improved glycemic control in patients with poorly controlled type 2 diabetes, without major adverse events (Signal Transduct Target Ther. 2025;10(1):371). While our current work focuses on defining the anatomical organization and functional asymmetry of this circuit in mice, the clinical findings suggest that the core principles, sympathetic control of hepatic glucose metabolism via CG-liver pathways, may be conserved and of translational relevance. Additionally, we clarified the interpretation of TH labeling and expanded the discussion of hepatic sensory and parasympathetic innervation, acknowledging their important roles in liver-brain communication and identifying them as key directions for future research. Collectively, these revisions provide a more balanced, clinically informed, and rigorous framework for interpreting our findings.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      We thank the reviewer for this suggestion. We agree that the species should be clearly indicated. The findings presented in this study were obtained in mice using tissue clearing and whole-organ imaging approaches. Due to technical limitations, these observations are currently restricted to the mouse strain. We have updated the title (Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice) and clarified the species used throughout the manuscript.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also hits a portion of sensory fibers that need to be ruled out in whole-mount imaging data

      We thank the reviewer for pointing this out. We acknowledge that TH labels not only sympathetic fibers but also a subset of sensory fibers. We have added a limitation of this point in the revised manuscript. In addition, using SyGlass (2.4.0) three-dimensional reconstruction, we observed TH-positive nerve fibers originating from the CG-SMG extending along the porta hepatis and penetrating into the liver parenchyma. Given that the CG-SMG is a well-established sympathetic ganglion innervating visceral organs (Nature. 2025 Jan;637(8047):895-902.), these nerve fibers can be definitively identified as sympathetic. In parallel, we collected DRG from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While T7-12 DRG are known to contain sensory neurons innervating the liver, only a sparse number of PRV-positive neurons were detected in these segments (Anat Rec A Discov Mol Cell Evol Biol. 2004 Sep;280(1):827-35. Auton Neurosci. 2024 Jun;253:103174). The additional figure and discussion content have been highlighted in the revised version as indicated (Discussion section, paragraph 6: “Third, although whole-mount TH immunostaining with three-dimensional reconstruction revealed sympathetic nerve bundles projecting from the CG to the liver, TH is not entirely specific and can also label a subset of sensory neurons. More selective approaches, such as genetic targeting of sympathetic lineages, will be important for further validation.”).

      Author response image 2.

      Representative immunofluorescence images of PRV-labeled neurons (EGFP) in DRG from the spinal segments T1-6 (bottom) and T7-12 (top) following PRV injections into the liver lobes. Scale bars, 200μm

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      We thank the reviewer for this suggestion. Previous studies largely relied on electrical stimulation to modulate liver innervation, which provides relatively coarse control of neural activity (Eur J Biochem. 1992;207(2):399-411). By contrast, our use of chemogenetic and optogenetic approaches allows selective, cell-type-specific manipulation of LPGi neurons. We revised the Discussion to place our functional data in the context of prior work, highlighting how these more precise approaches improve understanding of the contribution of liver-innervating neurons to hyperglycemia. The newly added discussion has been clearly labeled in the response to facilitate your review (Discussion section, paragraph 3: “This spatial organization is likely obscured by conventional electrical stimulation, which indiscriminately activates heterogeneous sympathetic fibers. By contrast, chemogenetic and optogenetic approaches permit selective, cell type-specific manipulation of LPGi neurons, thereby revealing the contralateral and lobe-specific architecture of brain-liver sympathetic control”).

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases to tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though it is clearly assumed to be. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We thank the reviewer for this insightful and important comment, which highlights a potential alternative interpretation of our findings. We agree that chemical sympathetic denervation with 6-OHDA may induce compensatory changes in non-sympathetic inputs, including sensory and parasympathetic (vagal) innervation of the liver.

      Conceptually, we agree with the reviewer’s perspective that the central nervous system operates as a highly integrated homeostatic regulatory system, continuously receiving and integrating a broad range of afferent signals. These inputs include, as noted by the reviewer, hepatic sensory and vagal afferents (Science. 2024;386(6722):673-677), as well as centrally derived interoceptive signals such as brain glucose, temperature sensing, even the pulsation of cerebral vascular system (Cell Metab. 2025;37(11):2264-2279.e10.; Cell Metab. 2022;34(6):888-901.e5; Science. 2024;383(6682):eadk8511). The CNS integrates these diverse signals and generates coordinated efferent outputs to maintain systemic homeostasis.

      From this viewpoint, the changes in c-FOS activity that we observe in the LPGi likely represent only a limited snapshot of this broader integrative process, rather than evidence of a single dominant pathway. We acknowledge that compensatory sensory or parasympathetic mechanisms, in addition to altered sympathetic drive, contributed to the observed LPGi activation following hepatic sympathetic denervation.

      We further acknowledge that, due to limitations in scope and experimental focus, we did not directly assess sensory or parasympathetic innervation of the liver in the present study. As appropriately pointed out by the reviewer, a more comprehensive characterization of hepatic neural inputs would provide a more complete picture of the underlying neurocircuitry. To address this, we expanded the Discussion and explicitly noted this limitation, including a more balanced discussion of potential crosstalk among sympathetic, sensory, and parasympathetic pathways and how these may collectively influence LPGi activity. For your convenience, the newly added discussion text has been distinctly marked in the manuscript (Discussion section, paragraph 4: “Although enhanced sympathetic output appears to mediate much of this compensation, our findings suggest that the underlying regulation extends beyond a purely descending pathway. In particular, c-FOS activation in the contralateral LPGi after unilateral 6-OHDA-mediated denervation suggests that the loss of peripheral input may be sensed through an ascending neural pathway, centrally integrated, and translated into compensatory sympathetic output to the intact hepatic lobes. These results therefore support a model in which hepatic glucose production is regulated by an integrated afferent-central-efferent loop, with our current analyses primarily resolving its efferent component.”).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Although the findings are interesting, this reviewer has major concerns about the experimental design, methodology, results, and interpretation of the data. Experimental details are lacking, including basic information (age, sex, strain of mice, procedures, magnification, etc.).

      We thank the reviewer for this important recommendation. We agree that comprehensive reporting of experimental details is essential for rigor and reproducibility.

      In the revised manuscript, we added complete information regarding mouse strain, sex, age, and sample size for each experiment. In addition, detailed descriptions of surgical procedures, viral constructs, injection parameters, imaging magnification, and analysis methods have been incorporated into the Methods section.

      These revisions ensured that all experiments are described with sufficient technical detail and clarity to allow accurate interpretation and replication of our findings. Experimental details have been incorporated, and corresponding revisions are highlighted for ease of review.

      Reviewer #3 (Recommendations for the authors):

      Addressing a few questions might help:

      (1) The study found that liver-associated LPGi neurons are predominantly GABAergic. It would be informative to molecularly characterize the PRV-traced, liver-projecting LPGi neurons to determine their neurochemical phenotypes.

      We thank the reviewer for this insightful suggestion. We agree that molecular characterization of liver-projecting LPGi neurons is important for understanding their functional identity.

      This issue has been addressed in detail in our recent study (Cell Metab. 2025;37(11):2264-2279.e10), in which we performed single-cell RNA sequencing on retrogradely traced LPGi neurons connected to the liver. These analyses demonstrated that the majority of liver-projecting LPGi neurons are GABAergic, with a defined transcriptional profile distinct from neighboring non–liver-related populations.

      Based on these findings, the current study selectively targeted GABAergic LPGi neurons using GAD1-Cre mice. We have explicitly cited these molecular results in the revised manuscript to clarify the neurochemical identity of the PRV-traced LPGi neurons. New text has been incorporated, and corresponding revisions are highlighted for ease of review.

      (2) Is it possible to do a local microinjection of a sodium channel blocker (e.g., lidocaine) or an adrenergic receptor antagonist into the porta hepatis? That would potentially provide additional evidence for the porta hepatis as the functional crossover point.

      We appreciated the reviewer’s thoughtful suggestion. Although pharmacological blockade at the porta hepatis can modulate local neural activity, this approach is inherently limited in its ability to distinguish between ipsilateral and contralateral inputs. Consequently, it may not provide definitive evidence for neural crossover at this specific site.

      In our view, the anatomical evidence provided by whole-mount tissue clearing, dual-labeled tracing, and direct visualization of decussating nerve bundles at the porta hepatis offers a more definitive demonstration of sympathetic crossover. Pharmacological blockade would affect both crossed and uncrossed fibers simultaneously and therefore would not specifically resolve the anatomical organization of this decussation.

      Nevertheless, we agree that functional interrogation of the porta hepatis represents an interesting direction for future work, and we acknowledge this possibility in the Discussion (Paragraph 6: “Fourth, although our data support a peripheral decussation at the porta hepatis, direct validation of this crossover site was not feasible with local pharmacological blockade, as currently available approaches lack sufficient spatial specificity and would likely perturb multiple neural components. Future studies employing more selective inhibitory strategies will be required to directly test this possibility.”).

      (3) It is possible to investigate the effects of unilateral LPGi manipulation or ablation of one side of CG/SMG on liver metabolism, such as hyperglycemia?

      We thank the reviewer for this important suggestion. Because unilateral LPGi manipulation was already examined in our study (Figure 2D), we focused here on unilateral ablation of the CG to further assess lateralized sympathetic control of hepatic metabolism. We successfully performed unilateral CG ablation without LPGi manipulation, but observed no significant change in blood glucose compared with the sham group (Author response image 3A and 3B). To determine whether glucose homeostasis was nonetheless affected, we further performed glucose tolerance tests (GTT) and insulin tolerance tests (ITT) (Author response image 3C and 3D). Neither test showed significant impairment after unilateral ablation, suggesting that compensatory neural mechanisms and/or hormonal homeostatic regulation may be recruited to preserve systemic glucose homeostasis.

      Author response image 3.

      (A) Blood glucose levels in mice subjected to left- or right-sided CG ablation via 6-OHDA treatment (n = 6). (B) Representative images of ablation of CG. Scale bars, 100 μm. (C and D) Blood glucose levels during GTT (C, n = 6) and ITT (D, n = 6) in mice with left- or right-sided CG ablation.

      Reviewer #4 (Recommendations for the authors):

      In the abstract and elsewhere, the use of the term 'sympathetic release' is unclear - do you mean release of nerve products, such as the neurotransmitter norepinephrine? This should be more clearly defined.

      We thank the reviewer for pointing out this ambiguity. We agree that the term “sympathetic release” was imprecise. In the revised manuscript, we explicitly referred to the release of sympathetic neurotransmitters, primarily norepinephrine, from postganglionic sympathetic fibers.

      We revised the wording throughout the manuscript to ensure accurate and consistent terminology and to avoid potential confusion regarding the underlying neurobiological mechanisms.

      Original: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased sympathetic release, glucose production, and glycogen depletion.”

      Revised: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased norepinephrine release, glucose production, and glycogen depletion.”

    1. eLife Assessment

      This valuable study demonstrates that self-motion strongly affects neural responses to visual stimuli, comparing humans moving through a virtual environment to passive viewing. The evidence for visuomotor mismatch responses is solid, although the interpretation in terms of prediction remains somewhat preliminary. This study bridges human and rodent studies on the role of prediction in sensory processing, and is therefore expected to be of interest to a large community of neuroscientists.

    2. Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference cannot be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features, but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course, the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course, this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information *from perception*. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Similarly, a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      Comments on latest version.

      Nice to see the added extra analyses. Can't see any more will be achieved via further rounds and happy with the summary to stand as is.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      - Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      - The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      - Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Comments on latest version.

      The authors added a brief discussion paragraph which addresses my previous comment.

    4. Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      Comments on latest version:

      The authors added useful points to the discussion and also included time frequency analyses to the paper formally, which strengthens the translational potential, in addition to the bolstering their claims slightly.

    5. Author response:

      The following is the authors’ response to the previous reviews

      We thank you for the time you took to review our work and for your feedback! The main changes to the manuscript are:

      We added a paragraph to the Discussion addressing differences in visuomotor mismatch responses recorded over frontal and occipital electrodes, and their possible interpretation.

      We added time-frequency power and phase-locking analysis as supplementary figures to the manuscript.

      We added a statement in the Discussion emphasizing the importance of performing these experiments with denser EEG channel coverage.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference can not be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active. I have a couple of questions/comments.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      The reviewer is correct. But one should keep in mind that our result here stands in the context of a considerable amount of work on mouse cortex investigating responses to very similar visuomotor mismatches. There we can we have much additional evidence to argue that the cortical response to a visuomotor mismatch is a prediction error. We would argue, it is the best one can do in human experiments.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      We have added a paragraph to the Discussion addressing the differences between signals recorded over frontal and occipital electrodes, as suggested.

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      We are not sure what the reviewer is referring to here. We certainly do not think (or are aware of having argued) that a visuomotor prediction error is the only possibly interpretation of the response. In the last paragraph of our response to the reviewers point 3 in the last revision, we explicitly discuss that the interpretation of the responses as a prediction error is only one possible interpretation. Our argument is that it is the most likely given the evidence.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      We would be happy to add additional considerations to other processes. If the reviewer has anything specific in mind, we can add that, but it would need to be somewhat concrete with some theoretical basis. We share the reviewer’s intuition, but unless this can be formalized to the point of being experimentally testable, we do not see any value in discussing it in the manuscript.

      Regarding the reason for a difference in visual responses in walking vs sitting state is, this is not entirely clear to us. Predictive processing would provide one possible explanation. Assuming the precision weighting of predictions is higher during walking, the sudden appearance of a visual stimulus might lead to stronger stimulus history prediction errors than when just sitting. But this is rather speculative.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information ‘from perception’. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Please excuse, the reviewer is correct, the orthodoxy of an idea is not relevant. For dubious reasons, we chose to euphemize what we actually meant to say here. With regards to circuit implementations of predictive processing (we cannot and do not intend to speak to interpretations of predictive processing that relate to conscious perception much of V1 activity is likely not consciously perceived – we assume this is what the reviewer is referring to by “we have one percept”) – the reviewers interpretation was not unorthodox, but rather incorrect (which is what we should have said). The statement that “the brain predictively ‘cancels’ expected action outcomes from perception” is incorrect in the context of sensory processing – based on both theoretical models of predictive processing, and more importantly physiological evidence. If the point was only in regards to conscious perception, we also suspect the statement is wrong, but even if it were correct, don’t see how it pertains to our work.

      Similarly a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      We believe it is absolutely essential for the progress of the field that we start to emphasize the differences between something that is “predictable in principle” and “predicted by the brain”. There is likely indeed a hierarchy of predictability that looks something like this:

      (1) Sensorimotor coupling

      (2) Laws of physics

      (3) Behavior of other living things

      (4) Artificial, human-made statistical relationships

      Almost all published experiments are based on the fourth type of prediction. Indeed, why not use physics simulations instead of oddballs and MMN? We absolutely should! But the field tends to revert to artificial couplings. As a direct consequence of this, the number of papers appearing recently (from both human and mouse fields), that are built on the following premise:

      (1) Expose an animal or human to an artificial coupling between A and B (e.g. an oddball, or a global oddball, or any of a myriad other constructions).

      (2) Probe for prediction error responses to the violation of the artificial coupling.

      (3) Find no prediction error responses and conclude predictive processing is wrong.

      Is utterly baffling. The fallacy here is of course the assumption that if something is predictable in principle, the brain must predict it. Thus, we are, and will continue to be strong on this point, and we think it is essential that we – as a field – are.

      Hope these comments are helpful.

      Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      We have added a discussion of this.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice, and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      The reference electrode placed at Cz makes it is difficult to interpret relative differences between frontal and occipital electrode responses, as the occipital electrodes are placed farther away from the Cz reference than the frontal electrodes. Similarly, signal occuring cortically near the Cz reference might only appear as though it is occipitally distributed in this montage. It is common in EEG research to remontage the data to an averaged common reference in order to better interpret the scalp distributions. As the electrode coverage was sparse for some subjects, this could be challenging, and this reviewer does not feel that it is necessary to do this analysis step, or even to drastically rewrite the body of the paper. We only request that some discussion, however brief, is included in the discussion section or the methods that recommend more dense electrode coverage in the future to better interpret scalp distributions and potential meso-scale sources.

      We have added a discussion of this as suggested.

      This is just a suggestion. The authors are encouraged to analyse (and report) time-frequency power and phase locking for these mismatch responses, as is common in much of the literature (see Roach et al 2008 Schizophrenia Bulletin). This is not to say that doing so will yield insights into oscillations per se, but converting the data to the time-frequency domain provides another perspective that has some advantages. fosters translations to rodent models, as ERP peaks do not map well between species, but e.g. delta-theta power does (see Lee et al 2018 Neuropsychopharmacology; Javitt et all 2018 Schizophrenia research; Gallimore et al 2023 Cereb Ctx). Further, ERP peaks can be influenced by the actual neuroanatomy of an individual (especially for quantifying V1 responses). Time frequency analyses may aid in interpreting the "early negative deflection with a peak latency of 48 ms " finding as well. As it stands, the report is complete, and it would be acceptable if the authors chose to save this type of analysis for a future publication.

      We have added this as suggested.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed most of my concerns by providing additional analyses, partly based on new data. The volume conduction issue is partly addressed based on the result showing latency differences, however to confidently assign responses to visual regions, one would need to perform recordings with a larger number of electrodes, sufficient to perform source localization. Nevertheless, the manuscript is now more solid than the previous version.

      We have now added this point to the Discussion.

      Reviewer #3 (Recommendations for the authors):

      The reviewer appreciates that the authors have carried out time-frequency analyses, and are ok with them leaving this out of this paper.

      We have now added this to the manuscript.

      Finally, in response to the participant quote "are you printing this? hi mom!" - this reviewer concedes that it does not significantly detract from the report, and, in the interest of amusement and joy, would abide its reinstatement.

      We greatly appreciate the reviewers entertaining our attempts at humor but will leave it out as originally suggested.

    1. eLife Assessment

      This study presents a key finding: self-generated mechanical stresses enable collective protocell proliferation without dedicated division machinery, offering insight into primitive life's population growth. While quantitative imaging, membrane tension measurements, and computational modeling support the mechanism, establishing causal links between deformation and division and testing sensitivity assumptions would strengthen the work. Overall, the work reports important findings, and although the evidence in support of the conclusions is largely solid, some incomplete elements need to be addressed.

    2. Reviewer #1 (Public review):

      Li and Wu, in this article, explore the proliferation of wall-less L-forms derived from Bacillus subtilis as mimics for protocells and report an interesting new mechanism for their proliferation. The authors carry out live-cell imaging of the L-forms and find that the clusters of cells forming proto-colonies proliferate better than the isolated single cells of L-forms. They further examine the causes for this indefinite proliferation of proto-colonies of L-forms, as compared to the isolated cells, which lyse and die out sooner. The authors show that when L-forms exist as isolated single cells, the growth in volume exceeds the rates at which surface area increases, leading to lysis. The authors further quantify the circularity and effective radius in growing proto-colonies, qualitatively estimate membrane tension and suggest that the confined space allows for mechanical shear in these cells. They propose that the mechanical stress on the membranes from adjacent cells in confined spaces deforms membranes and supports cell division to keep the population growing. These findings are also supported by modelling the proto-colonies in quasi-2D planes.

      The study is quite interesting and significant as it has implications for both evolutionary aspects as well as clinical importance, given the proliferation of certain pathogens as L-forms. The aspect of carrying out long-term imaging of colonies of L-forms as spatially constrained entities and the findings are fascinating. While the conclusions presented are backed by experiments, I only have a few questions concerning the proposed mechanism of division and proliferation of these proto-colonies.

      (1) The authors propose that the growth of neighbours leads to shearing forces in membranes and show that membrane tension increases at the periphery of the proto-colonies. They suggest that the increased membrane tension leads to a greater chance of deformation, enabling cell division. However, it is not quite clear how greater membrane tension could lead to cell division. Studies have suggested that membrane fluidisation is important for the cytokinesis event, which includes FtsZ-based division (Ramirez-Diaz, 2025).

      (2) Thus, it becomes quite important to rule out any role for the cytoskeletal proteins in the observed division with an increase in membrane tension. The authors note in line 188 that the division in protocells is independent of FtsZ, but this independence is for protocells that divide by extrusions and resolution, where the membrane is highly fluidised (Mercier et al., 2012).

      (3) The authors may use the L-form derivative where the FtsZ protein can be depleted and assess the proliferation of the proto-colonies. Likewise, authors should rule out the role of MreB as well.

      (4) Although the growth rates have been shown to be similar for proto-cells and the proto-colonies, and only the membrane tension has been shown to be higher at the periphery, it is also important that the authors rule out any increased lipid synthesis in the fraction of dividing cells in these proto-colonies. Without this, one could also envisage a model where membranes are fluidised due to an increase in lipid biosynthesis in a fraction of cells in these confined spaces, leading to increased vesiculations which experience membrane shear and deform. The authors can also consider examining proto-colonies of L-forms of branched-chain fatty acid-deficient strains.

      (5) Lastly, why does CellROX stain the proto-colonies? Are these tightly packed cells experiencing higher oxidative stress, and could that also contribute to membrane tension? This should at least be discussed.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Mechanical interaction enables a collective mode of protocell proliferation" addresses an interesting and potentially high-impact question about protocell proliferation in prebiotic environments. The central observation that wall-deficient B. Subtilis proliferate in dense colonies but die by membrane rupture in isolation is striking and a fundamental contribution to the field. However, the data and the mechanistic explanation offered for this observation are incomplete. The measurement and analyses used to build the mechanistic case raise methodological questions that may be difficult to fully resolve with the existing data and approach, and the authors should therefore consider whether additional independent experiments are needed to support the mechanical shearing hypothesis.

      Strengths:

      The central observation that wall-deficient B. Subtilis proliferate in dense colonies but die by membrane rupture in isolation is convincing and a significant contribution to the field interested in the growth of protocells. This adds an important aspect of collective growth that is different from individual dynamics.

      Weaknesses:

      (1) The surface-volume balance ratio η is an elegant concept and provides an intuitively reasonable framework for understanding why isolated cells lyse. However, its application here rests on treating cells as flat discs of uniform thickness, and Figure S4 makes clear that the cells are highly irregular and lobulated in ways that make this approximation questionable. The authors should clarify whether they have validated this assumption, for instance, through direct thickness measurements or sensitivity analysis. However, even with such validation, the modest quantitative differences between aggregated and isolated η trajectories, combined with the inherent difficulty of accurate perimeter measurement in these morphologically complex cells, mean that η measurements are unlikely to provide robust quantitative support for the mechanism. The authors should therefore consider whether η is better presented as a motivating conceptual framework rather than primary quantitative evidence and seek more direct experimental support for the surface-volume balance argument through independent means. For instance, osmotic pressure manipulation to test whether reducing volume expansion pressure preferentially rescues isolated cells.

      (2) The comparison of circularity between colony and isolated cells is complicated by the fact that the segmentation approach is fundamentally different in the two conditions; isolated cell boundaries are detected against a clear background, while colony boundaries are detected from inter-cell fluorescence gradients. The authors should address whether this introduces systematic bias. However, this may be difficult to fully resolve given the inherent complexity of the system, and that the deformation-division correlation in Figure 3C, while suggestive, would be substantially strengthened by a more direct perturbative approach. Specifically, can cell deformation be mechanically induced in isolated cells, for instance, using micromanipulation, external flow, or confinement in fabricated microstructures, to test whether artificially deformed isolated cells gain the ability to divide? Such an experiment would provide direct evidence for the deformation-division link that the correlational analysis cannot.

      (3) The interpretation of FliptR lifetime as a direct membrane tension readout is complicated in this system because cell-cell interfaces contain two apposed bilayers in proximity, potentially altering FliptR photophysics through changes in local membrane density and dielectric environment independently of tension. The authors should address whether they have considered this possibility and what controls were performed. Disambiguating tension-dependent from environment-dependent lifetime changes is technically challenging and suggests that the membrane tension argument would be more convincingly supported by an independent measurement approach. For instance, tether-pulling experiments using optical tweezers on isolated versus colony cell membranes, or testing whether membrane tension-modulating interventions such as osmotic shifts produce the predicted changes in cell fate, would provide more direct evidence. The current FLIM data should be regarded as suggestive rather than conclusive.

      (4) The Cellular Potts Model reproduces the experimental observations, but since its key parameters, particularly the substrate-pinning energy, were calibrated against those same observations, this demonstrates internal consistency rather than independent validation. The η-based lysis criterion is implemented as a model input, meaning the model cannot independently confirm the η hypothesis. The authors should clarify the extent to which model parameters were fitted to data versus independently motivated and be explicit that the model is best understood as a mechanistic illustration rather than independent evidence.

    4. Reviewer #3 (Public review):

      Summary

      This manuscript reports that protocells derived from wall-deficient B. subtilis proliferate well when densely packed but fail to divide and eventually lyse when isolated. The authors attribute this density-dependent proliferation to mechanical shearing between growing neighbors, which deforms cells and increases the likelihood of membrane stalk formation and subsequent scission, enabling division without any dedicated molecular machinery. Through a combination of quantitative imaging, membrane tension measurements, and Cellular Potts Model simulations, the authors make a compelling case that self-generated mechanical stresses are critical for sustaining population growth in protocolonies. The findings have implications for understanding the lifestyles of primitive life forms, L-form bacterial pathogenesis, and the design of synthetic cells.

      Strengths

      The central finding is both surprising and counterintuitive: crowding is not just tolerated by protocells but is required for sustained population growth. The mechanism the authors propose is interesting: mechanical shearing between growing neighbors deforms cells, increasing the likelihood of membrane stalk formation and thus division, all without dedicated molecular machinery. Conceptually, this is a type of biophysical "scaffold" (Jacobeen et al. 2018, Nat. Phys.; Day et al. 2022, Biophys. Rev.) in which key elements of a Darwinian loop, namely a life cycle involving growth and reproduction, are provided "for free" by physics, enabling open-ended Darwinian evolution that can eventually bring these life cycle components under developmental control. Such scaffolds, both biophysical and ecological (Black et al. 2020, Nat. Ecol. Evol.; Libby & Rainey 2013, Phys. Biol.), are likely key mechanisms in the origin of life and in evolutionary transitions in individuality, and this paper provides a nice example of how they can work in a protocell context.

      The combination of experiments and modeling works well. The membrane tension measurements are the strongest piece of evidence for the proposed mechanism, showing directly that tension is elevated in protocolonies and concentrated at cell-cell interfaces. The Cellular Potts Model captures the key experimental features. The discussion is nicely balanced, particularly the note about Gram-negative L-forms, whose rigid outer membrane may preclude this mechanism, which is a testable prediction for future work. I would suggest the authors also discuss the connection to biophysical scaffolding, as I think this is conceptually important and would help situate their work within a broader framework for understanding how primitive life cycles can arise from physical processes (see also Zamani-Dahaj et al. 2023, Genes; Hammerschmidt et al. 2014, Nature).

      Weaknesses

      The surface-volume balance analysis is central to the argument, and it depends on the assumption that cells have a fixed thickness of 0.8 µm, taken from the width of walled cells. But these are wall-deficient cells, which are mechanically quite different, and their thickness could plausibly vary during growth or under compression. I think the paper would benefit from either a direct measurement of cell thickness or a sensitivity analysis showing how η responds to plausible variation in this parameter. If the results are robust, that would put the analysis on much firmer ground.

      The positive correlation between cell shape deformation and division rate (Figure 3C) is central to the proposed mechanism, but I think the paper needs to be more careful about the jump from correlation to causation. The authors propose that deformation increases the likelihood of membrane stalk formation, leading to scission. That is plausible, but an alternative is that cells with higher local growth rates both deform more and divide more frequently, with the two outcomes driven independently by the same underlying cause. The paper does show that average volume growth rates are indistinguishable between aggregated and isolated cells, which argues against a simple "faster growth explains everything" interpretation, but this does not rule out local variation within protocolonies driving the correlation. I think the most convincing experiment would be to apply external mechanical stress to isolated cells and see if that alone can drive division, decoupling deformation from growth. I realize that this may be technically very difficult, but at a minimum, the paper should acknowledge this as an alternative hypothesis.

      The Cellular Potts Model has quite a few free parameters (Table S1), and it is not clear how tightly these are constrained by the data. A sensitivity analysis would go a long way toward showing that the results are robust and not overly dependent on specific parameter choices.

      In any case, this is a strong paper with a cool finding and an interesting mechanistic explanation. I think it will be of broad interest, particularly to people thinking about the origins of life and synthetic cell design.