10,000 Matching Annotations
  1. Jul 2026
    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATPbound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I mentioned that the final section of the Results part seems like an afterthought, especially since the heading suggests a broader scope.

      Reply: We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      The changes made to the section do not align with the wording of the reply. Please consider modifying it further.

      We appreciate the positive feedback and this final comment. We have revised the final section of the Results to better reflect the scope indicated by the heading. In addition to clarifying the kinetic analysis, we now explicitly relate our kinetic observations to previously determined thermodynamic measurements, showing that the rapid interconversion of ATP-bound conformations observed during steady-state turnover is consistent with a thermodynamic landscape characterized by a near-zero free-energy difference and entropy–enthalpy compensation. This revision more clearly integrates the kinetic and thermodynamic aspects of the transporter cycle.

    1. eLife Assessment

      This important study addresses a classic debate in visual processing, using a strong method applied to an impressive dataset obtained from a rare clinical population to evaluate hierarchical models of visual object perception. The paper provides compelling evidence that the hierarchical model is only partly supported: as expected, neural responses in ventral visual cortex show increased representational selectivity for faces along the posterior-anterior axes, but the onsets of the signals do not show a temporal hierarchy, indicating more parallel processing.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      Comments on revised version:

      The authors have performed additional analyses and checks and I would now rate the findings as important and compelling.

    3. Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The work is worth publishing for the data alone, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways, and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided, using resampling that preserve dependencies in a hierarchical manner, which is rare. Two types of analyses are also provided to support evidence in favour of the lack of practical significance for some of the comparisons. Collectively, the various analyses and rich illustrations provide convincing evidence in favour of the conclusions.

      Comments on revised version:

      The authors have addressed all my previous comments and the work is mostly limited by the lack of pre-registration and the exploratory nature of some of the analyses. However, with data and code available, other teams can assess the impact of researchers' degrees of freedom on the main outcomes.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      We would like to thank this reviewer for their positive comments and evaluation of our manuscript.

      Weaknesses:

      (1) A control analysis where areas have known differences in response onset should be performed to increase confidence that the proposed analyses would reveal expected results when a difference in response onset was present across areas. From Figure 3, it can be seen that many electrodes are placed in earlier visual areas (V1-V3) that have previously been shown to have earlier broadband responses to visual images compared to VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018). The same analyses as in Figures 4 and 5 should be used comparing VOTC to early visual areas to confirm that the analyses would detect that V1-V3 have earlier onsets compared to VOTC.

      First, we would like to mention that the analyses performed in our paper are commonly accepted analyses to extract time-domain information.

      Yet, the reviewer is right that, considering our claim, evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies would provide further support for our claims. The solution proposed by the reviewer is interesting but the number of face-selective recording contacts/sites in posterior occipital cortex (colored disks in figure 3) is too small for any meaningful comparison. Moreover, while absolute responses to visual images should indeed emerge earlier in early visual cortex than association cortex, it may not be the case for category-selective responses to visual images (which is what our claim is about).

      To address the reviewer’s concern, we used face-selective responses from regions known to have different response onset latencies: occipital and posterior temporal lobe vs. medial temporal lobe structures (e.g., Mormann et al., 2008, https://pmc.ncbi.nlm.nih.gov/articles/PMC2676868/). Waveforms and onset latencies for these regions (OCC, PTL, MTL) are shown in Author response image 1. Despite the small number of contacts showing significant face-selective activity in the MTL (N=20) and the lower SNR in this region, the onset latency differences between OCC/PTL and MTL are significant using all 4 methods of latency estimation (see methods in the main manuscript), with medium to large effect sizes. This was despite noisy latency estimates for the MTL (in particular, the ‘delta slope’ method could not be used to get meaningful Cohen’s d when comparing MTL to OCC). Latency estimates are also slightly higher for the ‘% of peak’ method than in the manuscript, as we estimated the latency at 25% of the peak (instead of 20% in the manuscript), again to allow meaningful estimations for the MTL.

      Author response image 1.

      In addition to this, we performed a simulation analysis where we statistically compared the measured PTL signals to ATL signals that have been artificially, incrementally, shifted forward in time. Author response image 2 shows (top row) the measured onset latencies differences between the 2 regions (PTL minus ATL) estimated using 4 different approaches as a function of the temporal shift applied to ATL, as well as the associated p-values (bottom row). As shown in Author response image 2, the ‘original’ unshifted data yields no significant difference between the 2 regions. The difference however becomes significant when ATL signals is shifted forward in time 10 to 30 ms, depending on the method used.

      Author response image 2.

      These 2 observations provide evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies, further supporting our claims.

      Last, as also suggested by reviewer 2, we conducted a thorough equivalence testing using ROPE and Bayesian factor to support the lack of differences between regions. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from early visual cortex than OCC). These analyses, now reported in the result section of the revised manuscript (Table 1), largely confirm the hypothesis of concurrent onset latencies across VOTC.

      (2) It is unclear why correlating mean timeseries helps understand how much variance is shared between regions (Figure 4). Any variance between images is lost when averaging time series across all images, and this metric thus overestimates the variance shared between areas. Moreover, the finding that correlating time domain signals across VOTC areas does not differ from correlating signals within an area could be driven by this averaging. For example, if the same analysis was done on electrodes in left and right V1 when half of the images had contrast in the left hemifield and the other half had contrast in the right hemifield, the average signals may correlate extremely well, while this correlation falls apart on a trial-by-trial basis. These analyses therefore need to be evaluated on a trial-by-trial basis.

      This is an interesting comment. We agree that variance between images is lost when averaging time-series across all images. However, to use the reviewer’s analogy, in order to support the claim that left and right V1 would show the same onset times and time-course (i.e., no hemispheric lateralization) for lateralized presentations (= the same kind of claim that we make in our paper), it’s the average response across images that should be compared, not a correlation run on a trial-by-trial basis (which would indeed falls apart because of a lack of response in the ipsilateral V1).

      Moreover, we would like to emphasize that the goal of this analysis in our paper is not to make claims about the variance shared between regions. In fact, this is not a key analysis in our paper, the outcome of which is not strictly necessary for the main argument made. Finally, if we were to perform a (time-consuming) image-by-image analysis in our study, correlations would be weak due to low signal-to-noise ratio (each face image appears only 1.6 times per stimulation sequence on average) and the fact that each face image appears after a different non-face image across presentations.

      (3) Previous studies on visual processing in VOTC have shown that evoked potentials are more predictive of the onset of visual stimuli than broadband activity (e.g. Miller et al., 2016, PLOS CB, https://doi.org/10.1371/journal.pcbi.1004660). Testing the prediction from a hierarchical representation that signals along the VOTC increase in onset time should therefore include an evaluation of evoked potential onsets in addition to broadband signals.

      We have used HFB responses in our study as these signals tend to be easier to characterize in the time domain than evoked potentials, and they are more local given their reduced SNR compared to evoked potentials (Jacques et al., 2022; https://pmc.ncbi.nlm.nih.gov/articles/PMC9457683/). Moreover we have previously shown highly correlated time courses across HFB and low frequency evoked potentials in the same paradigm (Jacques et al., 2022, eLife).

      Yet, to address this reviewer’s concern, we replicated the main analyses on low-frequency event-related potential signal, identifying contacts exhibiting significant face-selective responses in the same manner as in Jacques et al (2022). Namely, we start from bipolar-referenced sequences of recording corresponding to the full visual stimulation sequences (~70 s). For each recorded intracerebral contact, we average sequences in the time-domain, crop the average to contain an integer number of face frequency (1.2 Hz) cycles, run an FFT on the cropped sequences and identify the significant contacts with a Z-score procedure identical to that used for HFB signals. We then notch-filter out the visual response (6 Hz and harmonics) from the full length sequences, extract short epochs from the filtered sequences around the onset of each face image, average across epoch for each recording contact, subtract the mean signal measured in the baseline (-0.166 to 0 s relative to face onset) and take the absolute value (to be able to average across contacts despite differences of morphology and polarity). Significant contacts are then subjected to the same analyses as for the HFB signal.

      Results from these analyses are presented as supplementary material (Figure S9, Table S4) in the revised manuscript (referenced in lines of the main manuscript). While we were not able to obtain reliable latency estimates using the z-score method with the same parameters as for HFB signal, these analyses with ERP signal indicate similar onset latencies for ERPs as for HFB activity and largely replicate observations made with HFB. In particular, onset latencies were in a very similar range (~100 to 140 ms) with similar patterns across regions or along VOTC and between-region signal correlations. There were also a few significant face-selective responses over posterior ventro-medial occipital cortex, likely overlapping ‘early visual cortex’ (V1,V2v,V3v,hV4), probably due to limited low-level contributions in this paradigm (see Or et al., 2019, JOV; https://jov.arvojournals.org/article.aspx?articleid=2734585). Over these regions, onset latency was systematically earlier (up to 40ms) than in slightly more anterior regions, (i.e. anterior to -80 mm) where very little variability in onset latency was found up to the ATL region. We have acknowledged this in the revised manuscript.

      (4) Testing the second prediction, that the onset time of processing increases along the VOTC posterior to anterior path, is difficult using the iEEG broadband signal, because from a signal processing perspective, broadband signals are inherently temporally inaccurate, given that they are filtered. Any filtering in the signal introduces a certain level of temporal smoothing. The manuscript should clearly describe the level of temporal smoothing for the filter settings used.

      The reviewer is right that HFB signals are temporally smoothed, potentially yielding slightly underestimated onset latencies. However, our time-frequency analyses parameters ensured a minimal degree of smoothing. In fact, the original submission already contained a description of the expected temporal smoothing resulting from the wavelet transform. This is what we wrote in the original submission: “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had 20 ms of full width at half maximum (FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin), ensuring that onset timing information is accurate up to 10 ms (i.e half of the FWHM).”

      In the revised manuscript we further elaborate as follows:

      “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had a temporal spread of 20 ms (full width at half maximum - FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin). A simulation of HFB signals with a constant abrupt onset time and realistic signal-to-noise ratio indicates that the potential underestimation of onset latency due to the wavelet analysis is around 5-10 ms, which is on par with the value of the half width at half maximum (= FWHM/2 = 20/2 ms).”

      Author response image 3 displays simulated HFB signal (using identical wavelet parameters than in our manuscript) in an ideal scenario with a response starting at 150 ms in all trials (N=150 trials), reaching maximum 10 ms later. This provides a theoretical estimate of the slight underestimation of onset latency due to the wavelet transform. It shows onset latency estimates are at most 12 ms underestimation of true onset time.

      Author response image 3.

      That being said, given the physiological noise in the data, the fact that the response onset likely varies slightly from trial to trial, with a variable slope in activity increase, these wavelet parameters (within a certain margin) have likely little influence on the actual latency estimation.

      (5) The onsets of neural activity in VOTC are surprisingly early: around 80-100 ms. This is earlier than what has previously been reported. For example, the cited Quian Quiroga et al. (2023) found single neuron responses to have the earlier onset around 125 ms (their Figure 3). Similarly, the cited Jacques et al., 2016b and Kadipasaoglu et al., 2017 papers also observe broadband onsets in VOTC after 100 ms. Understanding the temporal smoothing in the broadband signal, as well as showing that typical evoked potentials have latencies compared to other work, would increase confidence that latencies are not underestimated due to factors in the analysis pipeline.

      In the revised manuscript, as suggested by reviewer 2, we have modified the data resampling strategy (using hierarchical bootstrap and permutation test that respects the nested structure of the data) to estimate onset latencies, confidence interval and permutation tests. Moreover, since the absolute onset latency estimates depend on the methods used, we now provide estimates using 4 different methods. The overall absolute onset latencies differ slightly across the 4 methods but all median onset latencies vary between 95 ms and 130 ms, which is similar to what was reported in some of the participants in Kadipasaoglu et al., 2017 (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0188834; note that latencies reported at the level of single sites or individual are usually higher due to lower signal-to-noise ratio). Moreover, these latencies for face-selective responses are actually very similar to those measured for non-selective/absolute responses to visual stimuli (e.g. Jacques et al., 2016: around 90-100 ms for faces in FG; Yoshor et al., 2007: ~100 ms in posterior Fusiform Gyrus; Regev et al. 2018: 90-120ms in posterior to middle FG). Other studies, measuring non-selective responses using more ‘conservative’ onset detection methods, find slightly later onset latencies (e.g. Cao et al. 2025: 139 ms in posterior FG; Martin et al., 2019: ~150 ms in ventral occipital -VO- regions).

      Cao R, Zhang J, Zheng J, Wang Y, Brunner P, Willie JT, Wang S. 2025. A neural computational framework for face processing in the human temporal lobe. Current Biology. DOI: https://doi.org/10.1016/j.cub.2025.02.063

      Martin AB, Yang X, Saalmann YB, Wang L, Shestyuk A, Lin JJ, Parvizi J, Knight RT, Kastner S. 2019. Temporal dynamics and response modulation across the human visual system in a spatial attention task: An ECoG study. Journal of Neuroscience 39:333–352. DOI: https://doi.org/10.1523/JNEUROSCI.1889-18.2018, PMID: 30459219

      Regev TI, Winawer J, Gerber EM, Knight RT, Deouell LY. 2018. Human posterior parietal cortex responds to visual stimuli as early as peristriate occipital cortex. European Journal of Neuroscience 48:3567–3582. DOI: https://doi.org/10.1111/ejn.14164, PMID: 30240547

      Yoshor D, Bosking WH, Ghose GM, Maunsell JHR. 2007. Receptive fields in human visual cortex mapped with surface electrodes. Cerebral Cortex 17:2293–2302. DOI: https://doi.org/10.1093/cercor/bhl138, PMID: 17172632

      As an important note, in the revised manuscript, we have removed data from 3 recording contacts from 1 participant that were located in the upper bank of the Calcarine Sulcus, which is actually outside of our VOTC region of interest. The 3 contacts being located in dorsal V1 or V2 were showing very early responses and were biasing our latency estimates for the OCC region.

      (6) Understanding the extent to which neural processing in the VOTC is hierarchical is essential for building models of vision that capture processing in the human brain, and the data provides novel insight into these processes.

      For additional context, a schematic figure of the hierarchical view and a more parallel system described in the paragraph on models of visual recognition (lines 553) would help the reader interpret and understand the implications of the paper.

      Our observations in the current study clearly indicates concurrent face-selective processing in the VOTC, which is incompatible with a serial hierarchical model. While we discuss how such concurrent activity could be implemented in the cortex (e.g., via direct input from ‘early visual cortex’ to different VOTC face-selective clusters), our data do not allow to provide more evidence in that respect to what already exists in the literature. Moreover, we are not providing data regarding connectivity (feedforward or re-entrant) either between face-selective regions or between these regions and ‘early visual cortex’.

      Author response image 4 shows a very simplified versions of standard hierarchical/serial versus concurrent/parallel models.

      Author response image 4.

      Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The data alone are valuable, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important, and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well-suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided. Collectively, the various analyses and rich illustrations provide superficially convincing evidence in favour of the conclusions.

      We thank the reviewer for their positive comments on our manuscript.

      Weaknesses:

      (1) The work was not pre-registered, and there is no sample size justification, whether for participants or trials/sequences. So a statistical reviewer should assess the sensitivity of the analyses to different approaches.

      Pre-registration of fundamental research in a clinical context is quite uncommon for intracranial data, owing, for instance to the time needed to accumulate data, or to the uncertainty of cortical sampling location in a given participant. Nevertheless, in the current study, sample size is much higher than in typical intracranial studies (usually 5-20 participants), in fact much higher than most typical Cognitive Neuroscience research. The same is true for the number of recording contacts (>10000 site here), and the number of trials considered for analysis. Each participant had a minimum of 164 face trials and an average of 262 trials (i.e. an average of 3.2 stimulation sequences of 82 trials), which is higher than most standard human electrophysiological studies.

      In the revised manuscript, to unsure that we have the maximum available power, and because our hypothesis is independent of hemisphere, we collapsed data across hemispheres for all analyses. We nevertheless provide analyses split by hemispheres as supplementary material.

      In addition, since onset latency estimations depend on the methods used, we now report onset latencies from 4 different methods (2 statistical and 2 non-statistical).

      (2) Frequentist NHST is used to claim lack of effects, which is inappropriate, see for instance:

      Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3

      Rouder, J. N., Morey, R. D., Verhagen, J., Province, J. M., & Wagenmakers, E.-J. (2016). Is There a Free Lunch in Inference? Topics in Cognitive Science, 8(3), 520-547. https://doi.org/10.1111/tops.12214

      Please see reply to the next comment (3).

      (3) In the frequentist realm, demonstrating similar effects between groups requires equivalence testing, with bounds (minimum effect sizes of interest) that should be pre-registered:

      Campbell, H., & Gustafson, P. (2024). The Bayes factor, HDI-ROPE, and frequentist equivalence tests can all be reverse engineered-Almost exactly-From one another: Reply to Linde et al. (2021). Psychological Methods, 29(3), 613-623. https://doi.org/10.1037/met0000507

      Riesthuis, P. (2024). Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science, 7(2), 25152459241240722. https://doi.org/10.1177/25152459241240722

      We thank the reviewer for pointing this out. In the revised manuscript we conduct and report a thorough examination of equivalence using ROPE and Bayesian factor to support the lack of differences between regions. We did not use TOST procedures as these have low power and require huge samples be meaningful (Riesthuis, 2024). Instead we relied on Bayes factor and descriptive proportion in ROPE. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from ‘early visual cortex’ than OCC). These analyses, now reported in the result section of the revised manuscript, along with effect sizes, confirm the hypothesis of concurrent onset latencies across VOTC.

      Riesthuis P. 2024. Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science 7.

      Detailed methods are reported as well:

      “In addition, to statistically assert whether onset latencies measured across main VOTC regions (OCC, PTL, ATL) were consistent with a concurrent (parallel) face-selective activation, we used Bayesian equivalence testing, relying on two separate metrics: (1) the percentage of differences in region of practical equivalence (ROPE), and (2) the Bayes factor using Cauchy prior. Equivalence bounds for ROPE were defined by combining two components: (1) a component of physiological variability and (2) a component reflecting expected delays in response onset between regions attributed to neural conduction delay, given the differential distances separating early visual cortex (EVC) from posterior face-selective regions (e.g. IOG) vs. anterior regions (ATL) and assuming signal mainly travels between VOTC regions through major postero-anterior axis fiber bundles of the Inferior longitudinal fasciculus (ILF) or the inferior fronto-occipital fasciculus (IFOF). Physiological variability corresponded to expected measurement noise and between-subject variability. Physiological equivalence bound was obtained by multiplying a Cohen’s d of 0.2 (conventional threshold for a negligible effect) with the (pooled) between-subjects variability in onset latency (computed for each region using a jackknife procedure and a correction factor of [N-participants – 1] to the jackknife standard deviation). For conduction delay bounds, the expected latency difference between regions under parallel activation depends on: (1) the distance between each region (considering the estimated origin of the signal is the same for all regions - EVC), and (2) neural conduction velocity. For inter-region distance we determined, for each region, the 5 and 95 percentile of the Talairach y-coordinate distribution and defined the maximal distance bounds as the distance between the y-coordinate corresponding to 5% of region 1 (e.g. OCC) to the coordinate corresponding 95% of region 2 (e.g. PTL). This resulted in the following maximum distance values: OCC-PTL: [54]mm; PTL-ATL: [54]mm; OCC-ATL: [83] mm. For conduction velocity, we used a constant value of 3.5 m/s, based on median axonal conduction velocity for cortico-cortical connections (Lemarachal et al., 2022; Van Blooijs et al., 2023), which was more conservative than using a range of values (e.g. 1.7 to 5.3 m/s based on Lemarachal et al., 2022). Maximum expected conduction delay was computed as [maximum distance / conduction speed] (e.g. for OCC to PTL: 0.054 / 3.5 = 15 ms). Resulting equivalence bounds (ROPE) were asymmetrical given than one region (e.g. OCC) is always closer to the source (EVC) than the other region (e.g. PTL) and were defined as [-1*physiologial_bound +1*physiologial_bound+max_conduction_delay]. For instance, using the z-score method to measured onset latencies, ROPE was [-12 to 27] ms for OCC to PTL, meaning that under equivalence, OCC can be activated up to 27 ms earlier than PTL (maximum conduction delay + noise), while allowing for some instances where OCC activates later (up to 12ms, due to noise only).

      For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (4) The lack of consideration for sample sizes, the lack of pre-registration, and the lack of a method to support the null (a cornerstone of this project to demonstrate equivalence onsets between areas), suggest that the work is exploratory. This is a strength: we need rich datasets to explore, test tools and generate new hypotheses. I strongly recommend embracing the exploration philosophy, and removing all inferential statistics: instead, provide even more detailed graphical representations (include onset distributions) and share the data immediately with all the pre-processing and analysis code.

      Data will be shared upon publication of the manuscript (see OSF repository in https://osf.io/2qzym). While we agree the dataset is large and could be explored in many ways, we do not consider the current study to be exploratory in nature. While our measurements could have turned out to clearly support hierarchical processing in human VOTC, our point in this manuscript is that the evidence derived from this large dataset unequivocally points instead toward concurrent activation of face-selective regions along the VOTC from IOG to antFG+ (i.e. a ~90 mm portion of cortex), with potential small variability accounted for by variability in axonal conduction velocity, signal-to-noise ratio or simple physiological variability. Other likely sources of variability such as type and density/size of fiber bundles across regions cannot easily be modeled with the current data set.

      (5) Even if the work was pre-registered, it would be very difficult to calculate p-values conditional on all the uncertainty around the number of participants, the number of contacts and the number of trials, as they are random variables, and sampling distributions of key inferences should be integrated over these unknown sources of variability. The difficulty of calculating/interpreting p-values that are conditional on so many pre-processing stages and sources of uncertainty is traditionally swept under the rug, but nevertheless well documented:

      Kruschke, J.K. (2013) Bayesian estimation supersedes the t test. J Exp Psychol Gen, 142, 573-603. https://pubmed.ncbi.nlm.nih.gov/22774788/

      Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin & Review, 14(5), 779-804. https://doi.org/10.3758/BF03194105 https://link.springer.com/article/10.3758/BF03194105

      All analyses and preprocessing stages are identical between regions and the number of trials is large enough not to be a constraining factor. As indicated above and below, we now report detailed equivalence testing and effect sizes and recomputed all statistics, taking into account the structure of the data as suggested by the reviewer.

      (6) Currently, there is no convincing evidence in the article to clearly support the main claims.

      Bootstrap confidence intervals were used to provide measures of uncertainty. However, the bootstrapping did not take the structure of the data into account, collapsing across important dependencies in that nested structure: participants > hemispheres > contacts > conditions > trials.

      Ignoring data dependencies and the uncertainty from trials could lead to a distorted CI. Sampling contacts with replacement is inappropriate because it breaks the structure of the data, mixing degrees of freedom across different levels of analysis. The key rule of the bootstrap is to follow the data acquisition process, and therefore, sampling participants with replacement should come first. In a hierarchical bootstrap, the process can be repeated at nested levels, so that for each resampled participant, then contacts are resampled (if treated as a random variable), then trials/sequences are resampled, keeping paired measurements together (hemispheres, and typically contacts in a standard EEG experiment with fixed montage). The same hierarchical resampling should be applied to all measurements and inferences to capture all sources of variability. Selectivity and timing should be quantified at each contact after resampling of trials/sequences before integrating across hemispheres and participants using appropriate and justified summary measures.

      The authors already recognise part of the problem, as they provide within-participant analyses. This is a very good step, inasmuch as it addresses the issue of mixing-up degrees of freedom across levels, but unfortunately these analyses are plagued with small sample sizes, making claims about the lack of differences even more problematic--classic lack of evidence == evidence of absence fallacy. In addition, there seem to be discrepancies between the mean and CI in some cases: 15 [-20, 20]; 8 [-24, 24].

      In light of the reviewer’s comment, we recomputed all timing analyses using a stratified hierarchical approach to evaluate confidence intervals (using bootstrapping), statistical comparisons (using permutation tests) and equivalence testing.

      This is what we wrote in the revised methods:

      “The first two timing parameters of face-selective response, onset and offset latencies, were quantified per main VOTC region using a hierarchical bootstrapping approach to respect the nested structure of the data (region > participants > contacts > trials). For each bootstrap iteration and each region, we first sampled participants with replacement. Within each sampled participant, we then sampled contacts and then trials within sampled contacts, with replacement. Resampled trials, then contacts within participants, then participants within a region, were successively averaged to obtain a bootstrapped region-level response from which we derived onset latency (4 different methods) and offset latency. We obtained bootstrap distributions of onsets/offsets using 2000 bootstrap iterations per region, allowing to compute the median and 95% confidence interval for these 2 parameters.”

      Then, later about permutation tests:

      “Statistical significance of latency differences between main VOTC regions was assessed using a hierarchical permutation test. We use a stratification approach to partition participants into a paired set (i.e. participants that had recording contacts in the two regions compared) and unpaired set (participants with contacts in a single region). For paired participants, the region labels were randomly swapped within subject (i.e., exchanging the signals from the two regions), thereby preserving all participant-, contact-, and trial-level structure while breaking the association between region and latency estimates. For unpaired participants, participants were randomly reassigned between regions while preserving the original group sizes, to generate pseudo-groups under the null hypothesis of no regional difference. In each permutation, signals were averaged across trials, then contacts, then participants within each permuted group and latency was computed and stored from the resulting region-level signals. We performed 10000 permutations to obtain a distribution of regional differences of latencies under the null hypothesis and determine the p-value as the fraction of the null distribution larger or smaller than the observed (non-permuted) difference.”

      And then about equivalence testing:

      “For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (7) Three other issues related to onsets:

      (a) FDR correction typically doesn't allow localisation claims, similarly to cluster inferences: Winkler, A. M., Taylor, P. A., Nichols, T. E., & Rorden, C. (2024). False Discovery Rate and Localizing Power (No. arXiv:2401.03554). arXiv. https://doi.org/10.48550/arXiv.2401.03554

      Rousselet, G. A. (2025). Using cluster-based permutation tests to estimate MEG/EEG onsets: How bad is it? European Journal of Neuroscience, 61(1), e16618. https://doi.org/10.1111/ejn.16618

      In fairness, we do not understand or share the reviewers’ concern here. Hundreds of fMRI or EEG studies use FDR or cluster tests to make inference about spatial or temporal location. We use FDR correction in one of the onset latency estimation method and only consider one-sided differences. Other methods in the revised manuscript do not use FDR correction.

      (b) Percentile bootstrap confidence intervals are inaccurate when applied to means. Alternatively, use a bootstrap-t method, or use the pb in conjunction with a robust measure of central tendency, such as a trimmed mean.

      Rousselet, G. A., Pernet, C. R., & Wilcox, R. R. (2021). The Percentile Bootstrap: A Primer With Step-by-Step Instructions in R. Advances in Methods and Practices in Psychological Science, 4(1), 2515245920911881.

      Again, we are not sure what the reviewer’s is referring to. The confidence intervals are computed on latency estimates from bootstrapped waveforms. In the revised manuscript, these waveforms are obtained by averaging (i.e. mean) resampled trials, resampled channels, resampled participants. A trimmed mean could not be applied in this condition, except perhaps when averaging across trials. But then the trimmed mean would have to be applied separately at each time sample which would disturbed within-, or between-trial, variability.

      (c) Defining onsets based on an arbitrary "at least 30 ms" rule is not recommended:

      Piai, V., Dahlslätt, K., & Maris, E. (2015). Statistically comparing EEG/MEG waveforms through successive significant univariate tests: How bad can it be? Psychophysiology, 52(3), 440-443. https://doi.org/10.1111/psyp.12335

      The rule of contiguous significant points is a heuristic that many researchers have used successfully to avoid spurious detection due to temporal autocorrelation. While we are aware that more sophisticated methods exist to correct for autocorrelation, such as cluster-based approaches, it is not directly usable since it requires comparing 2 conditions. The approach described in Piai et al., 2015 is interesting but incorrect as well since it relies on split-half simulations, which reduced signal-to-noise ratio, resulting in over estimated correction to be applied. In our revised manuscript, we rely on multiple methods to estimate onset latency, some of which not relying on this heuristic. Moreover, we apply plausible physiological constrains to our latency estimates, such as rejecting any onset before 40 ms after stimulus onset.

      (8) Figure 5 and matching analyses: There are much better tools than correlations to estimate connectivity and directionality. See for instance:

      Ince, R. A. A., Giordano, B. L., Kayser, C., Rousselet, G. A., Gross, J., & Schyns, P. G. (2017). A statistical framework for neuroimaging data analysis based on mutual information estimated via a Gaussian copula. Human Brain Mapping, 38(3), 1541-1573. https://doi.org/10.1002/hbm.23471

      (9) Pearson correlation is sensitive to other features of the data than an association, and is maximally sensitive to linear associations. Interpretation is difficult without seeing matching scatterplots and getting confirmation from alternative robust methods.

      We rely on Pearson correlation because this replicates the method used in Kadipasaoglu et al., 2017. It is also a widely accepted measure of (linear) relationship (in our situation we did expect linear or near linear relationships) in the literature. To address the reviewers concern, in the revised manuscript we nevertheless report, as supplementary material (Figure S10), the same functional connectivity analyses performed using the methodology and code provided in Ince et al. (2017). The results of this analyses are extremely similar to the results using Pearson’s coefficients.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 6, the response onset latencies are rendered in a smoothed manner on the brain surface. However, with this smoothing, variability between electrodes cannot be seen, and it would be better visualized in color, rendered in each electrode.

      Latency estimates computed at individual channels are noisy, which is why we do not report individual channels latencies but rather rely on averaging signals across contiguous channels, either across whole regions (Figure 4) or across smaller volumes as in Figure 6.

      (2) Onset latencies of 60 seconds seem extremely early compared to literature typically citing evoked responses with a latency of ~170ms. It would help if some additional sanity checks were shown, such as showing the latency of early visual responses. This would help with relative comparisons.

      In the revised manuscript the earliest median latency is 95 ms, which is in line with previous intracranial electrophysiology literature (e.g. Jacques et al., 2016; Jacques et al., 2022 ; https://pubmed.ncbi.nlm.nih.gov/26212070/; https://pubmed.ncbi.nlm.nih.gov/36074548/). The 99% confidence intervals can result in earlier latencies both due to some participants showing early responses and noise in latency estimates. Also please keep in mind that latency estimates are usually earlier when combining data across channels/participants compared to individual channels simply due to differences in SNR or across participants (see e.g. Kadipasaoglou et al., 2017).

      The reviewer indicates “…to literature typically citing evoked responses with a latency of ~170ms.”. We are assuming that they refer to the face-selective N170 ERP component measured on the scalp in EEG. Even with this ERP component, the face-selective response usually starts around 120-130 ms after stimulus onset (e.g. Rousselet et al., 2008; Jacques, Retter and Rossion, 2016; https://pubmed.ncbi.nlm.nih.gov/18831616/; https://pubmed.ncbi.nlm.nih.gov/27138205/) at scalp level. With the same highly sensitive paradigm as used here in EEG, we have systematically shown latency onsets of face-selective activity shortly after 100 ms (e.g., Retter et al., 2020; also Quek & Rossion, 2017) not accountable for by low-level visual cues (i.e., not present for phase-scrambled stimuli; Rossion et al., 2015; Or et al., 2019). Our latency onsets are also in line with spiking activity recorded with the same approach in the LatFG (Laurent et al., 2026) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5955677

      (3) Figure 6B shows that the variability of latencies in the ATL is larger than the variability of latencies in the PTL. It would be helpful to evaluate whether, rather than in mean onset latencies, there is a change in variability in onset latency along the VOTC.

      This is an interesting point. However, it is difficult to evaluate since SNR is reduced in the ATL compared to OCC or PTL (Jacques et al., 2022). As a result, any measured modulations in the variability of onset latencies along VOTC may simply reflect changes in the precision of latency estimation driven by SNR variability.

      (4) Line 415 typo: 'there appears to be no delay' instead of 'there appear to be no delay'.

      We thank the reviewer for their careful reading of our manuscript. This has been corrected.

      (5) The discussion states in lines 455-457 that "a large proportion of neuronal populations in anterior VOTC regions exhibiting similar activity to different face images independently of the context in which they appear". However, this claim about similar activity to different faces should be evaluated and tested at the single-trial level. In addition, it is not clear how context was varied in the experimental design.

      Context is variable because each face image appears directly after a different object image (or object images) in the sequence. We have shown also in previous studies with this paradigm in EEG that the time-course of face-selective responses is similar across base frequencies (3-15 Hz) unless the rate is too fast, and whether an orthogonal or explicit face categorization task is used (Retter et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32119982/). Note that we do not claim that activity is identical across images but similar – if it was not (largely) similar, the averaged response would be jittered and low.

      (6) Line 516-517: DTI does not provide evidence for whether connectivity is direct or not, and what the directionality of connectivity is between two areas. This sentence should therefore state "..., suggest independent connections between early visual cortex and face-selective regions...".

      This has been rephrased.

      (7) Line 555 in the discussion, the definition of low-level visual should be expanded to include other early visual areas that have been demonstrated to respond earlier than VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018), to avoid the suggestion that V1 directly projects synaptically to all of VOTC (e.g. Markov et al., 2014, Cerebral Cortex, https://doi.org/10.1093/cercor/bhs270).

      We are not proposing that V1 directly projects directly/synaptically to all of VOTC, i.e., without other low-level retinotoptic areas involved; only that face-selectivity in the association cortex is not organized hierarchically. We have revised this sentence.

      (8) Line 585, for the sentence: "with temporal synchrony strengthening their connections", evidence or citations should be provided.

      Citations have been provided.

      (9) It is not clear what is meant in the paragraph starting in line 571: do the authors suggest that top-down signals are not necessary for fast recognition of clear views of faces, or additionally argue that these top-down signals are not necessary for detecting ambiguous or degraded inputs as faces?

      Exactly: That top-down (i.e., descending) signals may contribute but would not be necessary for fast recognition of clear views of faces AND for detecting ambiguous or degraded inputs as faces.

      (10) No statement was provided on data or code availability.

      Data will be made available on a repository upon publication (see https://osf.io/2qzym).

      Reviewer #2 (Recommendations for the authors):

      (1) FDR correction: which one? Please provide a reference.

      We now provide a reference, both in the results and methods: Benjamini and Hochberg, 1995.

      (2) In the introduction, this statement is too strong: "arguably the most familiar and ecologically valid stimulus". It is unclear how static 2D representations of faces are the most familiar and valid stimuli. Could you rephrase this? What about other very familiar stimuli like letters, words and biological motion?

      This statement is not about static 2D images of faces, but faces in general (in their natural environment). We do consider human faces (in general, not restricted to laboratory context) to be indeed the most familiar and ecologically important stimulus, both from an ontogenetic and phylogenetic perspective, unlike written material.

      (3) About the questioning of a strict temporal hierarchy, this EEG reference comes to mind: Foxe, J. J., & Simpson, G. V. (2002). Flow of activation from V1 to the frontal cortex in humans. Experimental Brain Research, 142(1), 139-150. https://doi.org/10.1007/s00221-001-0906-7

      As confirmed by Foxe et al. ’s (2002) paper to which the reviewer is referring to, there is indeed ample evidence that areas in the dorsal stream or frontal cortex (e.g. FEF) are activated very soon after V1 and before many ventral stream regions (e.g. Lamme and Roelfsema, 2000; https://pubmed.ncbi.nlm.nih.gov/11074267/). While Foxe et al.’s 2002 is highly valuable, it can hardly be compared with our current study which looks specifically into ventral stream areas which are largely indistinguishable using scalp EEG as in Foxe et al. ’s paper.

      (4) Regarding statistical significance, there is no such thing as a "trend". The threshold for a trend should have been pre-registered and applied to both sides of the magical boundary, for instance, with matching conclusions for a "trend toward non-significance (p=0.04)". P values near 0.05 provide weak support against the null. I would suggest leaving it at that. Nothing special happens at 0.05.

      This no longer appears in the revised manuscript.

    1. eLife Assessment

      This important study provides convincing evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial. By combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia.

      Comments on revised version:

      Thank you for the opportunity to re-review this interesting paper.

      This reviewer struggled to follow along with the author's response letter. It was challenging because responses were brief and did not list the specific changes made, leaving the reviewer to search for changes in the text. As best I could tell, there were no changes made that matched some of the highest-priority suggestions.

      Critique #1 - This reviewer did not find the presentation of confidence intervals in the Abstract and other sections, which were suggested. Please note the format that was suggested in the original critique from Reviewer 1.

      Critique #2 - regarding removal of unsubstantiated claims and use of a p-value to compare analysis to a null hypothesis - it seems the authors skipped over this critique and did not address it.

      Regarding bias: the authors answered a different question than the one asked. The reviewer asked what proportion of transmissions were sampled; the authors stated that only communities from which phylo data was acquired were modeled. Was sampling 100% in those communities? Please provide the percentage and provide analysis that shed light on how sampling bias could impact the analysis.

      Regarding "cherries" - the reviewer did not understand the author's response. The query was regarding what percent of the total number of phylogenetic pairs (denominator) were the 355 that had high confidence in directionality (numerator). The response could be expressed be a proportion.

      The expectation of ART reducing the age of sources of transmission seems unrealistic to this reviewer. People on ART are not always adherent and can still transmit during gaps in adherence. ART dramatically increases life expectancy with HIV, which would have the opposite effect.

    3. Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex specific transmission dynamics in Zambia with a goal of allocation of resources.

      Strengths:

      Important analysis to hone in on key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches used, and while the phylogenetic data was markedly more limited, it mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses and providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce.

      Comments on revised version.

      The revised manuscript clarifies the impact and utility of this work and better allows the comparability of the two methods. Highlighting the differences (or lack thereof) between the undiagnosed and diagnosed population) simplifies the public health approach.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial; by combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength. However, the evidence supporting key conclusions is incomplete, with several claims insufficiently substantiated by the data presented. Improvements in data presentation (e.g., quantification of qualitative statements, statistical estimates, and clearer description of results) would substantially strengthen the paper.

      We thank the editor and reviewers for their positive comments. We have revised the manuscript in response to the points raised, as described below.

      First of all, we would like to summarise what we have changed regarding the presentation of summary statistics throughout the text. We agree that many of the statements in the original submission tended towards being qualitative. This was the result of shying away from presenting two separate estimates, with different ways of quantifying uncertainty, in the text. The phylogenetics could be presented as mean and confidence interval, while the IBM would need some measure of centrality (mean or median) and the highest density interval for a summary statistic (e.g. the mean age gap) as it varies over the posterior. These are not directly comparable. We have now changed this to present both where appropriate, with cautionary note about the difference between the CIs and HDIs (lines 257-260).

      We also were somewhat arbitrary regarding where we chose to summarise the posterior in the IBM or look at the best-fitting single simulation, and where we presented the mean as opposed to the median. We have done a considerable overhaul of what is presented in this revision:

      (1) We always present the posterior summary unless the level of detail is such that summarising uncertainty over the posterior is not feasible (e.g. in figures 3, 4 and 5). In the latter case we still use the best-fitting IBM replicate.

      (2) In the main text we always present the mean. For the phylogenetics the summary statistics are mean and confidence interval. For the IBM this is the posterior mean, and 95% HDI, of the mean of a particular statistic as calculated in each of the 1000 IBM replicates. For example, each replicate will have its own distribution of male source ages which have a mean value. These means also vary over the posterior, and a mean of them is calculated, as well as the HDI interval to represent posterior uncertainty. This “mean of means” may be a slightly confusing piece of terminology at first glance, but it allows us to properly capture posterior uncertainty in a way we mostly avoided in the first submission.

      One result of 1) above is a change to figure 6. It is now summarised over the posterior, with the result that time trends that were previously not evident become clear. This changes our conclusions slightly (lines 529-537) but it should be noted that the magnitudes of the trends remain small.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia. The current manuscript draft is methodologically strong, but needs revision to strengthen the take-home messages. As written, there are many possible take-away conclusions. For example, the agreement between IBM and phylogenetic analysis is noteworthy and provides a methodological focus. The revealed age patterns of transmission could be a focus. The effects of the PopART intervention and the consequences of a 1-year disruption could be a focus. It is important, though, that any main messages summarized by the authors are substantiated by the evidence provided and do not extrapolate beyond the data that have been generated. I recommend that the authors think deeply about what the most important, well-supported messages are and reframe the discussion and abstract accordingly.

      We have rewritten the abstract, and also made changes to the discussion in order to centre our message around the contribution of particular of demographic groups to transmission, and how, with that contribution revealed, such groups can be selected for specialised interventions.

      Strengths/weaknesses by section:

      (1) ABSTRACT

      The Abstract summarizes qualitative findings nicely, but the authors should incorporate quantitative results for all of the qualitative findings statements.

      The abstract in the revision is extensively revised, and contains quantitative estimates throughout, from both methodologies where appropriate.

      The ending claim is not substantiated by the modeling scenarios that have been run: "targeted interventions for demographic groups such as under-35 men may be the key to finally ending HIV." It is straightforward to run this specific scenario in the model to determine whether or not this is true.

      Our modelling framework is not set up to model the “last mile” of HIV elimination, notably as it has no component for MSM or FSW transmission, and we do not feel that we could confidently present results regarding it. As a result, this statement has been greatly softened in the new abstract (lines 75-78).

      The authors should add confidence intervals to the quantitative metrics, such as the 93.8% and 62.1% incidence reduction.

      These have been added.

      (2) RESULTS

      The authors should check the Results section for any qualitative claims not substantiated by the analyses performed, and ensure the corresponding analyses are presented to support the claims.

      The Results and Methods describe the model's implementation of the PopART intervention differently. The Methods describes it as including VMMC, TB, and STI services, while the Results only mentions intensified HIV testing and linkage.

      This is a slight misreading of the text. That paragraph in the Methods is describing the trial itself, not the modelling framework.

      A limitation of the model is that HIV disease progression is based on the ATHENA cohort in the Netherlands, which is a different HIV subtype (B) than the one in the research setting (C). The model should be configured using subtype C progression data, which have been published, or at least a sensitivity analysis should be conducted with respect to disease progression assumptions.

      The available literature does not suggest a significant difference in progression between subtypes B and C, and we have added text and citations to this effect (lines 699-701).

      In Table 2, the authors should consider adding a p-value to establish whether or not IBM and phylogenetics estimates are different.

      We have done this; the appropriate test was a posterior predictive check. See lines 261-263, 575-579 and 805-814.

      (3) DISCUSSION

      The literature review and comparison of study results to previously published phylogenetic studies is very nice. The authors could strengthen this by providing quantitative estimates with CIs for a more scientific comparison of the study results vs. prior studies, perhaps as a table or figure.

      We have expanded the discussion on this point (lines 504-527). We considered adding a table, but the existing literature that directly answers the questions we ask is quite limited and fragmentary. For example, Monod et al do not present a complete treatment of age gaps. The literature using regression analyses to identify predictors of HIV prevalence or incidence related to partner age is extensive, but those results are not directly comparable to ours.

      The authors state that due to "the narrow geographical catchment area... The results should not be automatically extrapolated to apply to other SSA settings." The authors should exercise this caution when comparing the results to studies in South Africa and elsewhere.

      We have made more explicit acknowledgements of these limitations (lines 598-600).

      There are many other limitations to the analysis, including some mentioned above, that are not acknowledged. The authors should think carefully about what the most important limitations are and acknowledge them honestly at the end of the Discussion section.

      The limitations paragraph has been revised (lines 598-605).

      Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex-specific heterosexual HIV transmission dynamics in Zambia, with the goal of allocating resources.

      Strengths:

      Important analysis to hone in on the key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches were used, and while the phylogenetic data were markedly more limited, they mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses. The authors did a nice job of providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce

      Weaknesses:

      To increase the impact and utility of this work, it would be helpful to parse the analysis just a bit further to estimate the roles of undiagnosed vs diagnosed and untreated subpopulations on this transmission. PopART is a multifaceted intervention, but the cost, effort, and approach to reengagement in care vs testing/treatment can be quite different.

      We have now provided stratified results by diagnosed and non-diagnosed status of the source, as well as an overall summary of the proportion of undiagnosed sources by age and sex. See lines 305-310, 539-547, and table 3.

      Recommendations for the authors:

      Reviewing Editor:

      We commend you for conducting a rigorous and comprehensive study titled "The age and sex dynamics of heterosexual HIV transmission in Zambia: an HPTN 071 (PopART) phylogenetic and modelling study" that significantly advances the understanding of HIV transmission dynamics in sub-Saharan Africa. The study utilizes an innovative dual-methodology approach integrating individual-based mathematical modelling (IBM) and pathogen phylogenetics to characterize heterosexual HIV transmission patterns by age and sex during the PopART trial in Zambia.

      This manuscript reports on HIV transmission dynamics in Zambia using data from the PopART study, combining individual-based modelling and phylogenetic analysis. The use of two independent methodologies enhances confidence in the consistency of the findings and enables robust cross-validation. The work addresses an important topic in HIV prevention, particularly in settings where resources may become more constrained, and offers insight into potential demographic targets for intervention.

      However, several aspects of the manuscript limit its current impact. The main take-home messages are diffuse and not clearly presented. Some conclusions in the abstract and discussion appear to go beyond the scope of the presented data. For instance, the claim that targeting under-35 men may be key to ending HIV is not directly tested in the modelling scenarios and should be reframed or removed unless supported by new analyses. Furthermore, important quantitative details, such as confidence intervals, p-values, and precise age group estimates, are lacking in key sections (e.g., the Abstract and Results).

      The authors are encouraged to clearly identify and communicate their central findings, ensure all claims are fully supported by their analyses, and make the data more accessible to readers by adding detailed, quantitative summaries where needed.

      The following are our recommendations to the Authors:

      (1) Clarify Study Objectives and Central Messages

      Reframe the abstract and discussion to highlight a clear, well-supported set of main findings.

      Avoid overgeneralized or unsubstantiated claims, especially those not directly tested by your model (e.g., the effectiveness of targeting under-35 men).

      As stated above, we have revised this text accordingly.

      (2) Support Qualitative Claims with Quantitative Data

      Provide numerical results, including effect sizes and confidence intervals, wherever qualitative trends are mentioned.

      For example, restate: "The largest gaps for female recipients were among the youngest" as "... in the age group XX-YY with OR = Z.Z (95% CI: A.A-B. B)."

      As mentioned at the top of the review, we have overhauled the treatment of summary statistics extensively, and now give confidence or highest density intervals throughout the text.

      (3) Improve the Results Section

      Check that all claims are supported by the analyses, and ensure figure references are accurate.

      The statements that went beyond what was supported, notably about ending the epidemic by targeting young men, have been removed. The typo in table references has been fixed.

      Annotate Figure 6 with trendline coefficients and p-values where applicable.

      The takeaway message of figure 6 has now changed and we no longer see no trend, just a minor one.

      Revise Figure 4 for clarity or consider replacing it with a tabular format.

      We would prefer to keep the current figure 4, as we have not found any clearer way to illustrate the patterns, which are the consequence of the phenomenon observed in figure 5. We have put more explicit descriptive text in the discussion, linking the two figures (lines 470-476).

      (4) Address Potential Bias and Model Assumptions More Rigorously

      Explain sampling bias in IBM and phylogenetics (e.g., how the 355 high-confidence phylogenetic pairs were selected).

      The reviewer comment regarding the 355 pairs was based on a misapprehension; we used all the pairs we found using the phyloscanner pipeline. There are no sampling bias issues involved in the IBM as every individual in the simulations is considered. Appendix 2 includes some sensitivity analysis results if the procedure used to find the 355 is changed.

      Discuss how the use of subtype B disease progression data from the ATHENA cohort may impact results in a subtype C setting. A sensitivity analysis would strengthen this.

      Subtype B progression data was used in the absence of any appropriate data from subtype C, but the literature does not suggest any major difference between the two (lines 699-701).

      (5) Include More Detail on Undiagnosed Populations and ART Effects

      Estimate the roles of undiagnosed and untreated subpopulations in driving transmission.

      As mentioned above, this analysis has been added.

      Clarify mechanistically how ART might influence age gaps in transmission dynamics.

      This now is clarified in the introduction (lines 127-129).

      (6) General Improvements

      Provide p-values where comparisons are made (e.g., in Table 2).

      Use consistent terminology and definitions across Methods and Results.

      Add more discussion on limitations, especially regarding generalizability to other SSA settings.

      All of these have been inserted as previously mentioned.

      By addressing these points, the manuscript would present a more coherent narrative and a stronger, evidence-based contribution to the field. We appreciate you all for your fantastic effort and hope you will reflect the feedback in your final paper.

      Reviewer #1 (Recommendations for the authors):

      Thank you for the opportunity to review this interesting manuscript.

      In the public review, I have recommended that the authors should incorporate quantitative results for all of the qualitative findings statements. As one example, I would recommend that "We found the largest gaps for female recipients were among the youngest of those recipients" is re-written as "The largest gaps for female recipients were in the age group XXX-YYY with OR=ZZZ (XXX-YYY)." such as odds ratios, and specific outcome definitions including ages. To give one more example: "immediate increase in the average age at transmission of both sources and recipients" could be rephrased as "increase in the average age at transmission by XXX (YYY-ZZZ) years for sources and XXX (YYY-ZZZ) for recipients over [TIME PERIOD]."

      We hope the revisions we have made to the statistical presentation are satisfactory as a response to this request.

      Again in the public review, I recommended checking the Results section for any qualitative claims not substantiated by the analyses performed, and ensuring the corresponding analyses are presented to support the claims. An example is: "Trends are minor or non-existent in the former two variables." - please annotate Figure 6 (assuming the authors meant to reference Figure 6 and not 7 here?) to show over what period trendlines were fit and provide the coefficient and CI. To support the stated claim even more strongly, a p-value might be apt with a null hypothesis of a slope of zero.

      Please check the numbering on all figure references in the text, as some appear to be misnumbered. E.g., where the text refers to Figure 7, I believe the authors meant to reference Figure 6.

      The change to how we handled the statistics has changed the message of figure 6 (which is now figure 7) and rendered this somewhat moot. We have checked that all figure and table references are now correct.

      Figure 3 is very nice, but if the axes were flipped on one panel, it would make them easier to compare, and then adding some statistics to assess whether the patterns are the same or different when a man vs woman is the source.

      We have flipped the axes here.

      Figure 4 was too complicated for me. I could not follow the Sankey flows because there is too much going on and overlapping. Consider revising to make it easier to digest... perhaps to table format?

      As mentioned above, we would prefer to keep this figure, but we have situated it better in the text.

      Reviewer #2 (Recommendations for the authors):

      A few points that would improve the clarity and the strength of the manuscript

      (1) There is a need to clarify more about how the IBM and phylogenetic data does not suffer from sampling bias. For e.g.,

      Line 205: What proportion of the transmissions modeled in the IBM from Zambia?

      All of them. We confined the analysis of the IBM to the Zambian communities from which phylogenetic data was acquired (lines 755-758).

      Line 217: What proportion of the phylogenetic pairs (cherries) suggesting transmission were the 355 that had high confidence in directionality. How do these pairs compare to the others

      There was no identification of “cherries” involved in picking these pairs; the phyloscanner procedure does not use that step. We confined our analysis solely to the pairs for which we did identify a direction of transmission; that is the 355. Appendix 2 includes a sensitivity analysis involving varying the parameters by which these were identified.

      (2) I appreciate the authors noting that MSM transmissions are unlikely to be playing a role in this cohort, as noted in previous work by the group. However, systematic undersampling of men is common in other study cohorts of HIV. While the MSM and heterosexual networks may be relatively distinct, undersampled men who are bridging the networks could impact the estimates. Can the authors use the time to diagnosis analysis (HIV phyloTSI) to estimate rates of undiagnosed men and women?

      We feel that this is beyond the scope of this work. The phylogenetics dataset in its totality could be used for this purpose (although it is probably highly biased towards undiagnosed individuals due to the considerable majority of samples coming from the healthcare facilities). However, we concentrate here solely on the subset involved in our probable transmission pairs, which is fairly small. Extending the scope to an exploration of the full dataset would seem like a separate study, which we do have plans to do.

      We have used the IBM for this question instead (lines 303-321), however, as MSM transmission was not modelled, it is also not ideal for answering this question. Ultimately we feel that the way these studies were implemented makes it an unsatisfactory tool for answering the MSM question, important as it is.

      (3) Expanding on the point above, in other settings, transmission to young men has been associated with partnerships with older men, and if these young men then transmitted to young women, would we see a similar effect as noted in these models (assuming the young men were less well sampled).

      Our previous work (Hall et al., 2024) suggested no excess of identified male-male pairs in the phylogenetics dataset which might suggest cryptic male-to-male transmission. The age disparities would be worth exploring had this been found, but is curtailed by the lack of it.

      (4) Related to the point above, is there an estimate of the populations (age and sex) that are undiagnosed in the IBM model? Can this be teased out... is transmission from men to women more likely 2/2 lack of diagnosis... or lack of engagement in care?

      We have explored results by diagnostic status as it pertains to age and sex, but we feel that moving on to a more general exploration of the role of diagnosis and lack of engagement in care is again going beyond the scope of what is already a long paper.

      (5) I'm still not fully clear as to why ART might affect age gaps. Can this be explained in more detail?

      See lines 127-129.

    1. eLife Assessment

      This valuable study describes PXGS, a poly-transgene expression system that exploits the mutually exclusive splicing of Dscam variable exon 4 to enable conditional, simultaneous expression of up to 12 transgenes in Drosophila, addressing a longstanding limitation in which conditional co-expression has been restricted to a handful of genes. The approach is conceptually elegant and technically accessible, with potential applications spanning neuroscience, synthetic biology, and biomanufacturing across arthropod species. The evidence that Dscam exon 4 splicing is preserved in a UAS vector and that individual alternates can be replaced with functional transgenes is solid, and the in vivo axonal re-wiring application provides a convincing proof of principle. Quantitative characterization of expression levels, a direct demonstration of expression across all twelve positions, and additional imaging controls would further substantiate the system's utility and scope.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes the development of an expression system enabling up to 12 transgenes using the alternatively spliced fourth exon of *Drosophila* *Dscam* gene under the control of a UAS element. This will be a useful tool if expression is needed in *Drosophila* cells (in culture or in vivo). where *Dscam* splicing machinery is active, which limits its use.

      Strengths:

      The tool developed is based on a well-established genomic element. The underlying idea is relatively simple yet effective.

      Weaknesses:

      The authors describe the weaknesses of their system well, most importantly, depending on the presence of adequate levels of Dscam splicing factors in targeted cells. This likely limits effective use of the methodology to some cell lines (e.g., S2) and certain tissues (nervous system and innate immune system). The manuscript could do a better job in showing protein expression levels more quantitatively, either in comparison to other methods or as absolute values (transcript numbers, protein molarity, etc.).

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Yu et al seek to develop a Drosophila genetic tool to simultaneously co-express up to 12 transgenes. They leverage the native Dscam exon 4 alternative splicing to generate a UAS to enable cell- and temporal-specific expression of transgenes. This tool is called the poly-transgene expression system (PXGS). Previous approaches to co-express transgenes have been limited to four to five genes, so PXGS would be a significant advancement, especially when examining processes that require robust expression of many genes to confer function. The authors showed that PXGS can drive expression of multiple (1) fluorescent reporters and (2) cell surface receptors in different cell types (neurons, glia, and muscles). However, there are major proof-of-principle experiments missing to demonstrate the utility of PXGS and its potential limitations. Additionally, some of the data is just not interpretable, and experimental rigor is significantly lacking.

      Strengths:

      Developing a genetic tool to co-express transgenes beyond what is currently available would be significant.

      Weaknesses:

      (1) While the authors stated that each PXGS construct can express 12 transgenes, this was not directly tested - the largest number of genes tested was in the PXGS_fluorophores, which has 4 genes inserted in 10 alternates (and therefore it can be determined if genes from all the alternates are spliced in at a meaningful level).

      a. First, the authors state that they tested the expression of the fluorophores in S2 cells using RT-PCR before generating the fly line. However, this data is not shown. Also, it is possible to test expression and localization of UAS transgenes in S2 cells with a ubiquitous GAL4, similar to what they did for GFP expression in Supplemental Figure 1.

      b. In the corresponding figure for this experiment (Figure 2), there are some concerning expression patterns and potential channel bleed-through/crosstalk. The nSyb-Gal4 is a pan-neuronal driver, yet expression of three fluorophores was extremely minimal. This could potentially be explained by the deterministic vs random alternative splicing. However, it is more concerning that in the GFP, RFP, and iRFP channels, the exact same tiny cluster of neurons is observed, suggesting potential bleed-through of the channels. Appropriate controls are required, including expression of a traditional fluorescent reporter with the nSyb-Gal4 they are using. And replicates would help with the experimental rigor.

      c. Additionally, it would be helpful to know if genes inserted at each alternate exon are expressed at a similar efficiency (vs. some alternate exons have higher levels of expression)

      (2) PXGS expression in non-neuronal cells: The authors attempt to show that fluorophores targeted to different cellular compartments can be expressed in neurons and non-neuronal cells (glia and ubiquitously).

      a. In Supplemental Figure 2A, they first use S2 cells to confirm expression, which does confirm. However, they use no markers to show that the fluorophores localize to the corresponding compartments (e.g., mitochondria and nucleus). In Supplementary Figure 2B, with that magnification and resolution, it is impossible to determine if the fluorophores localize properly.

      b. In Figure 3, it is impossible to know if there is any glial expression based on those images. They state that "subcellular localization of fluorophores was observed in the flight muscle", but again, that cannot be concluded from the images. Also, the schematic of the construct is the same one used in Supplementary Figure 2. Why show it again?

      (3) In the functional expression of PXGS transgenes section, while the authors used RT-PCR to show that each receptor gene is transcribed in S2 cells, it was not tested if they are correctly expressed, translated, and localized in the fly. The functional outcome observed (Figure 4 b-c) could be the result of misexpression of one or multiple genes.

      a. Supp Figure 3: Why does the Sli lane have so many bands?

      b. Why were these genes chosen for misexpression? Is there any evidence that they are required for the wiring of the mechanosensory neuron? Does the co-expression lead to an additive effect?

      c. The RT-PCR result (Supplemental Figure 3) showed variations (e.g., kek and kir being significantly dimmer than tutl; multiple products for Sli). Is there any explanation behind this, and could this be the outcome of some alternate exons being more efficiently spliced than others?

      d. Figure 4: These images seem to be taken with a widefield scope and only one plane. Is it possible that some of the pSC neurons are in a different Z plane, and they are not being captured here? There definitely is part of the axon terminal out of focus in some of the images. Also, most of the figure graph axes (e.g., 4b) are extremely difficult to read. And the figure overall is not easy to interpret.

      (4) Supplemental Figure 5: This figure is quickly mentioned in the Discussion without much explanation. First, this must be in the Results section since it is an experiment. Second, this needs more context because, as is, it seems like it was just thrown into the manuscript.

      (5) The authors mentioned that the size of the inserted genes could be a limitation for this technique and tested cell surface receptors of different sizes. However, there was no explicit discussion in the main text.

    4. Reviewer #3 (Public review):

      This paper, by Brian Chen and collaborators, adapts the highly alternatively spliced Dscam1 gene locus for use in a system for simultaneous multi-transgene expression in a variety of insect species. Specifically, they show that the hypervariable Dscam1 exon 4 region maintains its alternative splicing when placed in a UAS expression vector, and that each of the twelve exon 4 alternates can be replaced with an exogenous gene such that co-expression of up to twelve proteins can be achieved. Since the co-expression of more than a few proteins simultaneously is difficult, this represents a significant advance with multiple use cases. The authors validated the technique by assessing expression in vitro and in vivo, and by rewiring Drosophila sensory neuron axons by simultaneously expressing several cell surface receptors within the neuron. Overall, this is a clearly written paper that describes a potentially important new system. I have no major criticisms.

    5. Reviewer #4 (Public review):

      From the Reviewing Editor:

      All three reviewers recognized the conceptual originality of PXGS and its value to the Drosophila community and the broader multi-gene expression field. The core demonstration - that Dscam exon 4 mutually exclusive splicing is maintained in an exogenous UAS vector, and that individual exon alternates can be replaced with genes of interest for conditional in vivo expression - was viewed as solid and creative. The in vivo application re-wiring pSc axonal arbors using PXGS constructs loaded with cell surface receptors was noted as an encouraging functional validation.

      The reviewers differed in their overall enthusiasm. Reviewer 3 found the experiments straightforward, the results clear, and the system a significant advance with broad use cases, with no major criticisms. Reviewer 1 viewed the evidence as broadly solid, with the principal limitation being a lack of quantitative expression data and a dependence on adequate Dscam splicing-factor levels that constrain the system's applicable cell types - a limitation the authors themselves describe well. Reviewer 2 was the most critical, finding the underlying concept significant but the supporting data insufficiently rigorous to conclusively establish the tool's utility. The points below reflect the areas where reviewers - principally Reviewers 1 and 2 - felt the manuscript could be strengthened, should the authors choose to revise.

      (1) Quantitative characterization of expression. Reviewers 1 and 2 both noted the absence of a quantitative comparison of PXGS-driven expression - in absolute terms (transcript numbers, protein amounts) or relative to standard UAS constructs. Given that signal is inherently divided across 12 alternates per transcription event, characterizing expression efficiency and whether all alternates are spliced and expressed at comparable levels would substantially strengthen the manuscript. Reviewer 2 specifically asked whether genes at different exon 4 positions are expressed with similar efficiency.

      (2) Direct demonstration of 12-transgene expression. Reviewer 2 noted that, although the manuscript claims expression of up to 12 transgenes, this was not directly tested - the largest construct placed 4 distinct genes across 10 alternates, and the largest functional test used 3 genes per construct. Either a direct demonstration with more positions occupied or a more carefully bounded claim supported by the probabilistic framework would address this.

      (3) Controls and interpretability of fluorophore expression. Reviewer 2 raised concerns about Figure 2, where expression of three fluorophores under nSyb-Gal4 was minimal, and the same small neuronal cluster appeared across the GFP, RFP, and iRFP channels - raising the possibility of channel bleed-through. Appropriate controls (including a conventional single UAS-fluorophore driven by the same nSyb-Gal4) and replicates were requested. Reviewer 2 also noted that the S2 cell validation data for the fluorophore constructs, described as having been performed prior to fly line generation, are not shown.

      (4) Non-neuronal expression and subcellular localization. Reviewer 2 noted that the compartment-specific localization claims (mitochondria, nucleus) in Supplemental Figure 2 and Figure 3 are not supported by co-markers, and that the magnification and resolution in key panels are insufficient to confirm proper localization, particularly in glia.

      (5) Functional expression of receptor constructs. Reviewer 2 noted that, while RT-PCR confirms transcription of each receptor in S2 cells, correct translation and localization in the fly were not directly tested, so the observed phenotypes could reflect mis-expression of one or a subset of the genes. Clarification of why the specific genes were chosen, whether co-expression produces additive effects, and the cause of the variable RT-PCR band patterns (e.g., the multiple Sli products, dimmer kek and kir signals) was requested.

      (6) Supplemental Figure 5 (synthetic biology / RNAi). Reviewer 2 noted that this figure is mentioned only briefly in the Discussion despite representing an experiment, and recommended moving it to the Results with appropriate context. Reviewer 1 separately queried the meaning of the two white boxes in this figure and whether the RT-PCR convincingly supports expression of all genes shown.

      (7) Figure quality and labeling. Both reviewers flagged that the labels in Figure 4 (particularly panel d) and the axes in Figure 4b are too small to read. Additional labeling points were noted (alignment of "Repo-GAL4" and "brain" in Figure 3; unlabeled images in Supplemental Figures 2 and 4; clarification of whether the two lanes per group in Figure 1c are replicates or use different primers). Reviewer 1 also suggested that some figure legends describe conclusions rather than what is shown, and recommended that legends describe the data with interpretation kept to the text.

      (8) Additional points. Reviewer 1 suggested showing more of the gel in Figure 1 to demonstrate the absence of non-spliced fragments; clarifying the 1/12 probability argument (or moving it to the Discussion with transcript-number context); providing sequences for the fluorophore variants and fusion tags; and minor prose corrections ("Regardless if" → "Regardless of whether"; "dependent on three things" → "dependent on three factors"). Reviewer 2 noted that the gene-size limitation, though tested, is not explicitly discussed in the main text.

    6. Author response:

      We are pleased that the reviewers viewed the core demonstration (that Dscam mutually exclusive splicing is preserved in a vector and that exon alternates can be replaced with genes of interest) as a solid foundation for the system. We agree that the manuscript would be strengthened by clearer quantitative characterization of expression, additional controls for fluorophore imaging, improved presentation of the figures, and more precise wording about the current scope of evidence. In a revised manuscript, we plan to address these points by adding or clarifying quantitative expression analyses, including S2 cell validation data, adding appropriate imaging controls where available, revising claims about 12-transgene expression to distinguish design capacity from direct experimental demonstration, and improving figure labels and legends throughout.

      We also plan to expand the discussion of PXGS limitations, including cell-type dependence on Dscam splicing machinery, possible position effects, and gene size considerations. Finally, we will improve Methods reporting by adding resource identifiers, cell culture quality control information, statistical design details, and data/code availability statements where appropriate.

      We appreciate the opportunity to revise the manuscript and believe these changes will make the strengths and limitations of PXGS clearer to readers.

    1. eLife Assessment

      This study presents a valuable finding on linking the frequency of neural activity to cortical depths of blood flow in a naturalistic setting of participants listening to music. The presentation of evidence in the version of the original submission is incomplete, as further clarifications in methods and results, as well as performing additional analyses, would strengthen the study. The work will be of interest to cognitive neuroscientists working on multimodal recordings, auditory perception and music.

    2. Reviewer #1 (Public review):

      Summary:

      In their submitted work, Lee and colleagues examine the correlation between electrophysiological activity as measured by SEEG, and layer-specific activation patterns, as measured through 7T fMRI. This analysis was performed using patients undergoing monitoring for epilepsy surgery guidance, as well as healthy controls, as they both listened to music.

      They find that, in general, higher-frequency SEEG activity correlated positively with the fMRI signal, while lower frequencies correlated negatively. Across cortical depth, higher-frequency activity correlated positively with middle-to-upper layers, whereas lower frequencies showed their strongest negative correlations in superficial layers.

      Strengths:

      This is an interesting physiological study in that, to the best of this reviewer's knowledge, it has not been done before with auditory stimuli using the combination of iEEG (as opposed to scalp EEG) and fMRI. The framework fits well with models of layer-specific feedforward versus feedback processing (e.g., Bastos et al., 2012).

      Weaknesses:

      Its main limitations are a lack of specificity to the acoustic stimuli, the absence of correction for venous draining, and the fact that it is largely a replication/port of prior work.

    3. Reviewer #2 (Public review):

      Summary:

      The authors present an investigation of the relationship between the iEEG frequency bands signal and hemodynamic responses at different cortical depths. Based on this, the authors aim to uncover the layered origin of iEEG signals at different frequencies. The authors then interpret their results in terms of feedforward and feedback processing, arguing that the correlations between fMRI and iEEG signals reflect the interaction between both processes. In addition, the authors aim to infer the extent to which these processes are involved during naturalistic music processing.

      Strengths:

      This study combines the neural recording methodologies yielding the highest spatio-temporal precision achievable in humans, while using naturalistic auditory stimuli. This combination of recording methods and experimental design offers key insights regarding the precise origin of iEEG signals, which is necessary to improve the interpretability of future iEEG studies.

      Weaknesses:

      (1) The current framing of the paper leads the authors to interpret their findings in ways that are not warranted by the data. The main analysis of the paper consists of correlating the hemodynamic responses from different layers with iEEG signals from different frequency bands, which enables us to infer the relationship between the two signals. It does not, however, enable us to draw inferences regarding the extent of feedforward and feedback processing and the interaction between the two during naturalistic auditory processing. This would require comparing hemodynamic responses in different cortical layers or frequency bands activation against some baseline condition. Based on the presented analysis, statements such as "our frequency-specific results demonstrate that naturalistic music perception seamlessly integrates both feedforward and feedback processing streams" should be removed.

      (2) The presentation of existing literature omits key details and findings, making it difficult to fully understand the research question the authors are trying to address. For example, the author mentions studies showing that feed-forward and feedback processing are segregated across cortical layers and that feed-forward and feedback processing have distinct time-frequency signatures (lines 47-58). However, the authors do not mention which cortical layer or which frequency band is associated with which kind of processing. As a result, it is difficult for the reader to determine what exact hypothesis the author is trying to test in the study. This might also relate to the confusion raised in (1).

      (3) The method section omits key details. When describing the paradigm, the authors do not describe how the tones were presented, nor how the signals were synchronized. Similarly, there is no mention of the pipeline used for iEEG electrodes localization. In addition, the exact regressors that entered the generalized linear model of hemodynamic responses are not clearly stated: were all regressors (frequency bands + acoustic signal + HFA) entered together in a single model or in separate models? The mention of a cubic spline is also not sufficient for the reader to understand what was done and for which purpose. Finally, the exact tests used for some comparisons are omitted (in Figure 2, for example, no mention of the exact test used to compare betas between A1 and A2). The current structure of the method section is also quite difficult to follow: the authors switch back and forth between describing acquisition protocols and participant counts, for example.

      (4) The lack of methodological details (as described in point 3 above) casts doubts about the validity of some of the statistical tests reported. Throughout the paper, the authors present quantitative statements and statistical tests comparing the fitted beta parameters between brain regions (A1 and A2) and cortical depths. However, the authors do not mention any normalization procedure taken to ensure that the scale of the signals being compared was equated. If the overall magnitude of the signals in A1 differs from that of A2, the mean of the beta distribution is expected to differ as well. Similarly, if the signal-to-noise ratio differs between brain regions or cortical layers, so should the variance of the beta parameters across subjects, which might break the homoscedasticity assumption of some tests, which might or might not be a problem depending on the exact test the authors used (hence the importance of reporting them).

    1. eLife Assessment

      This study provides a useful anatomical resource by mapping the expression of four putative chemoreceptors in spinal cerebrospinal fluid-contacting neurons (CSF-cNs) of larval zebrafish. These descriptive findings offer an interesting entry point to explore how the nervous system senses signals within the spinal fluid microenvironment. The evidence supporting the spatial expression patterns of these receptors is convincing, utilizing high-resolution hybridization chain reaction (HCR) to validate previous transcriptomic data. However, the evidence remains incomplete regarding the actual functional roles of these receptors, as the study lacks protein-level validation, evidence of ligand availability in the CSF, or functional assays to demonstrate active chemoreception.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript examines the expression of putative chemoreceptors in CSF-contacting neurons of the larval zebrafish spinal cord. Using in situ hybridization, the authors show that sstr2a is preferentially expressed in ventral CSF-cNs, whereas grm2a, ptprna, and ldlrad2 are detected in both ventral and dorsolateral CSF-cNs, with additional expression in neighboring cells around the central canal.

      Strengths:

      The study provides useful anatomical information on the expression of putative chemoreceptors in CSF-contacting neurons. The experiments appear to be carefully performed, and the results are clearly presented with high-quality illustrations and informative schematics.

      Weaknesses:

      This work remains largely descriptive and based on mRNA expression. Therefore, the proposed roles in chemoreception, ligand sensing, lipid capture, or long-range CSF signaling remain speculative without protein-level or functional validation.

    3. Reviewer #2 (Public review):

      Summary:

      Verran et al. leverage a previously published RNAseq dataset of zebrafish cerebrospinal fluid contacting neurons (CSF-cNs) to identify potential receptors involved in chemosensory signalling in these neurons. They then validate expression of the identified receptors by hybridization chain reaction (HCR) in zebrafish larvae. This way they uncover potential roles for the somatostatin receptor Sstr2a, metabotropic glutamtate receptor Grm2a, LDL receptor Ldlrad2 and the Phosphatase receptor Ptprna, suggesting the existence of numerous chemo-sensory pathways in CSF-cNs and providing a potential entry point for further investigation.

      Strengths:

      This is a useful resource; the provided HCR data that demonstrates expression of these receptors in CSF-cNs is convincing, and the finding that CSF-cNs express these receptors is interesting.

      Weaknesses:

      The overall insight provided by this manuscript is rather limited, essentially just demonstrating the expression of 4 receptors in CSF-cNs, whose expression was predicted to be enriched in these neurons anyway by a previously published dataset.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to identify new molecular pathways that could enable long-range signaling through the cerebrospinal fluid (CSF), focusing on a specialized class of neurons called CSF-contacting neurons (CSF-cNs) in larval zebrafish.

      Strengths:

      Anatomical validation of transcriptomic candidates using HCR, providing high-resolution spatial mapping of chemoreceptor expression in CSF-contacting neurons and neighboring spinal cord cells. The work broadens the potential understanding of CSF-cNs and offers a resource for future functional investigations of CSF-mediated signaling.

      Weaknesses:

      The principal limitation of the study is that the conclusions remain largely transcriptomic and inferential. Although HCR convincingly validates mRNA expression, no protein-level evidence is provided to demonstrate receptor translation or subcellular localization, leaving uncertainty regarding functional receptor availability at the CSF interface. Moreover, the study does not establish whether the proposed ligands are present in the relevant CSF microenvironment or engage the identified receptors in vivo. As such, the functional significance of the proposed chemosensory pathways remains speculative.

    1. eLife Assessment

      This work provides a reassessment of VBIT-4, a compound previously proposed to inhibit oligomerization of the crucial protein known as the mitochondrial voltage-dependent anion channel. Combining complementary experimental approaches with molecular dynamics simulations, the authors provide compelling evidence that VBIT-4 primarily disrupts lipid membranes and induces channel-independent cytotoxicity. The study has important implications for interpreting previous work using VBIT-4 as a probe of channel function and highlights the need to consider membrane-disruptive effects when evaluating drug mechanisms.

    2. Reviewer #1 (Public review):

      Summary:

      The Voltage-Dependent Anion Channel 1 (VDAC1) is the most abundant β-barrel protein in the outer mitochondrial membrane and the main conduit for metabolite and ion exchange between the cytosol and mitochondria. Its oligomerization has been proposed to control mitochondrion-mediated apoptosis, making it a prime target for therapeutic intervention in diseases associated with excessive cell death, such as neurodegenerative disorders and autoimmunity. VBIT-4 is a small molecule developed to inhibit VDAC oligomerization and has shown therapeutic potential in various preclinical models. Despite its widespread use, the mechanism of action of VBIT-4 has not yet been fully elucidated. In this paper, Ravishankar et al. combine a suite of biophysical approaches with computer simulations to demonstrate that VBIT-4 forms water-permeable defects in membrane bilayers without any detectable effects on VDAC1 channel properties or oligomerization. Furthermore, cytotoxicity assays revealed identical VBIT-4 IC50 values in wild-type and VDAC1-KO cells, indicating that its activity does not depend on VDAC1. Collectively, these findings cast significant doubt on the widely held assumption that VBIT-4 is a specific inhibitor of VDAC1 oligomerization. Instead, it appears that VBIT-4 functions as a membrane-active compound.

      Strengths:

      This is a carefully conducted and well-written study that highlights potential side effects of VBIT-4, a compound that has been used to study the role of VDAC1 in a range of physiological and pathological conditions. The work is of interest to a broad readership by showcasing the importance of a systematic assessment of drug-membrane interactions to identify potential off-target membrane-driven effects of small molecules that may be mistakenly attributed to the inhibition of specific proteins. Its strength lies in the variety of complementary approaches the authors used to rigorously challenge the effect of VBIT-4 on VDAC1 organization and function. Overall, the experimental data are compelling and of high quality.

      Weaknesses:

      The authors used high-speed atomic force microscopy (HS-AFM) to study the impact of VBIT-4 on VDAC1 oligomerization in real time at nanoscale resolution. Toward this end, they adsorbed POPC:POPE:cholesterol membranes reconstituted with or without VDAC1 on mica. This revealed that the addition of VBIT-4 produced small perforations in the bilayer that were independent of VDAC1. In the absence of VBIT-4, VDAC1 showed the characteristic honeycomb topography that the authors described in a previous study (Reference 17). To quantitatively assess whether VBIT-4 affects VDAC1 organization, they analyzed protein compaction within clusters using inter-protein distance measurements. This analysis revealed no significant difference in VDAC1 organization between control conditions, 1 uM and 10 uM VBIT-4, supporting a model in which VBIT-4 primarily perturbs the lipid matrix rather than VDAC1 assemblies. This conclusion is based on the assumption that VDAC channels retain some lateral mobility in bilayers adsorbed onto mica. Do the authors have evidence that this is indeed the case? Did they also perform HS-AFM on VDAC1-containing membranes treated with VBIT-4 prior to adsorption onto mica?

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript challenges the widely used interpretation of VBIT-4 as a specific inhibitor of VDAC1 oligomerization, arguing instead that it acts primarily as a membrane-active compound. Using high-speed atomic force microscopy, electrophysiology, liposome leakage assays, Laurdan fluorescence, microscale thermophoresis, coarse-grained molecular dynamics, and cell-based assays in wild-type and VDAC1-knockout HeLa cells, the authors show that VBIT-4 partitions into lipid bilayers, induces membrane defects and leakage, and causes VDAC1-independent cytotoxicity within a concentration range commonly used to infer VDAC1-specific effects.

      Strengths:

      The main strength is the convergence of several independent approaches to the same conclusion. Atomic force microscopy directly visualizes VBIT-4-induced defects in lipid regions while VDAC1 assemblies remain apparently intact. Electrophysiology separates VDAC1 channel behavior from background membrane conductance and shows that VBIT-4 does not measurably alter VDAC1 conductance or voltage gating, while increasing nonspecific membrane permeability. Lipid-only membranes, lipid nanodiscs lacking VDAC1, and VDAC1-knockout cells provide important controls supporting a VDAC1-independent mechanism.

      The wild-type versus VDAC1-knockout cytotoxicity comparison is a particularly strong test of VDAC1 independence. The observations that VBIT-4 is poorly soluble, aggregation-prone, and sensitive to storage conditions are also important, as they offer a plausible explanation for variability across previous studies. The revised manuscript is further strengthened by quantitative analysis of VDAC1 organization in atomic force microscopy images and by simulations including multiple VDAC1 molecules.

      Weaknesses

      The main limitation is that the conclusion that VBIT-4 does not affect VDAC1 oligomerization is strongest for the specific readouts used here: atomic force microscopy measurements of cluster compaction, VDAC1 channel properties, and simulated assembly behavior. These are direct and informative measurements, but they are not identical to the chemical cross-linking readouts used in much of the prior VBIT-4 literature. Readers should therefore distinguish between VDAC1 cluster organization in membranes, as measured here, and cross-linking-defined VDAC1 proximity.

      A second limitation is the uncertainty around effective VBIT-4 concentration. Because VBIT-4 is poorly soluble, aggregation-prone, pH-dependent, membrane-partitioning, and storage-sensitive, nominal added concentration may differ substantially from the concentration of active compound available in each assay. This complicates comparisons across the different in vitro, simulation, cellular, and previously published assays.

      The coarse-grained simulations provide a coherent mechanistic framework for membrane partitioning, aggregation, and defect formation. However, the VBIT-4 coarse-grained model is newly parameterized and is used to support a quantitative partitioning argument. The manuscript would be easier to interpret if the coarse-grained-derived partition coefficient were reported with uncertainty, convergence information, and protonation state, and compared with a matched all-atom octanol-water partition estimate from the same atomistic model used to build the coarse-grained mapping. This matters because the partitioning argument is used quantitatively to relate micromolar aqueous VBIT-4 to millimolar concentrations in the bilayer.

      Finally, the cellular data strongly support VDAC1-independent cytotoxicity, but the lower-dose mitochondrial functional phenotypes were not directly compared between wild-type and VDAC1-knockout backgrounds. VDAC1 independence is therefore more directly established for cytotoxicity than for the lower-dose mitochondrial phenotypes.

      Overall, this work provides a valuable and timely reassessment of VBIT-4, and its central conclusion will be useful for researchers interpreting studies that use this compound as a probe of VDAC1 function.

    1. eLife Assessment

      This valuable study investigates the mechanisms underlying inter-item biases in visual working memory. By experimentally manipulating the relative noise levels of target and non-target items, the authors report bias patterns that are broadly consistent with predictions of their previously proposed normative demixing theory. However, the supporting evidence remains incomplete, as the manuscript lacks a sufficient description of the underlying theory, key assumptions, and a quantitative link between the model and behavioral data. The manuscript would be substantially strengthened by clearer exposition and stronger tests, including analyses of the full error distributions and comparisons with alternative models, which would increase its potential interest to the cognitive neuroscience and computational cognitive science communities.

    2. Reviewer #1 (Public review):

      Summary:

      Many previous studies have reported inter-item biases in visual working memory tasks. These biases can be either attractive or repulsive, depending on the particular experiments. It has been difficult to explain these biases in a unifying theoretical framework. Recently, Chetverikov (the first author of the current manuscript) proposed a demixing model for explaining these biases in Ref 22. That paper shows that both attractive and repulsive biases could emerge in the demixing framework depending on the noise properties. The current manuscript seeks to test the predictions of the demixing model experimentally in a series of new experiments and find evidence supporting the demixing model.

      Because previous modeling results described in reference 22 (which is a preprint) are essential in interpreting the results reported in the current manuscript, I also studied that preprint and used the results reported in that paper to help interpret the results in this paper. My comments below will also contain discussions of that modeling paper.

      Strengths:

      Overall, the computational model tested in the paper is novel and interesting.

      The demixing framework represents an appealing hypothesis that deserves further investigation.

      The current paper provides new empirical data showing that the target stimuli with the same absolute noise level can be either repelled from or attracted to non-target items, depending on the relative noise levels. The observation that biases depend on the relative noise levels is by itself an interesting one, and is consistent with the prediction of the demixing model.

      Weaknesses:

      While this manuscript contains interesting new experimental observations and theoretical ideas, it has several substantial problems in its current form, which limit the conclusions that can be drawn. The description of the computational model is too brief. The key modeling assumptions need to be better motivated and explained. As the computational models generate different predictions in different regimes, it is a bit difficult to evaluate how well the experimental data support the model at a more quantitative level. Also, the results focused on studying the biases in the behavior; it is unclear whether the model can fully explain the behavior data (such as error distributions or behavioral precision).

      Major concerns:

      (1) Concerns/suggestions regarding the computational modeling

      The current paper seeks to test the predictions of the demixing-based computational model proposed in reference 22. There are several problems with the modeling component in the current paper.

      (1a) The description of the model is too brief and difficult to understand. Although the model was proposed in reference 22, it would still be beneficial to provide more details of the model so that readers can understand and appreciate the strengths/limitations of the model.

      The generative model and the inference procedure could be better explained to better link the model to the behavior. In particular, how was the observer's behavioral report in each trial modeled? This requires more explanation because currently the demixing procedure estimates four parameters for a given trial, yet for a given trial, only one behavioral report was produced (e.g., current Experiment 1), or two reports were produced sequentially (e.g., current Experiment 2).

      (1b) Key modeling assumptions need better justification.

      One such key assumption is that on a given trial, each stimulus triggers many samples (or approximately, an entire response distribution), rather than a single sample. This assumption deviates substantially from prior work on ideal observer models. It was not clear whether this assumption is realistic. For the type of stimuli used in the current experiments, perhaps one can argue that each pixel corresponds to one sample of brain activity, thus collectively each stimulus should trigger many samples of activity in the brain. If this were to be the case, it would have two implications. First, the noise parameter in the model should be directly related to the magnitude of the stimulus noise. Thus, one should be able to plug these experimentally-controlled parameter values into the model to directly generate predictions about the biases. Second, when using stimuli with no stimulus variability (e.g., simple grating stimuli), the predicted biases should change. However, it wasn't clear whether this would hold experimentally, i.e., using gratings would lead to different biases or no biases.

      If the variability of the samples for a given stimulus involves neural noise, it would be useful to justify why it is reasonable to consider that many samples were generated per stimulus.

      (1c) As mentioned in (1b), the model assumes that on each trial, a large number of samples was generated. It would be useful to study and report how the prediction would change when the number of samples generated per stimulus is small. In particular, what happens when each stimulus only generates one measurement? This might be useful for interpreting previous experiment results with grating stimuli.

      (1d) Reference 22 studies how the predicted biases depend on the d-prime of the identifying dimension and found that the pattern of the biases varies substantially depending on the information available for the identifying dimension. However, the current paper didn't really discuss this important point. It is also unclear what parameters the authors used for the d-prime of the identifying dimension. Was it fitted directly to the data? The Methods section has some description on the "identifiability dimension", but it was a bit obscure.

      Intuitively, when the d-prime of the identifying dimension is very large, the demixing problem becomes irrelevant. In this case, there should not be any biases induced by demixing. In the case of the d-prime for the identifying dimension is 0, the problem should reduce to the simplified 1-d problem studied in reference 22. If my reading of reference 22 was correct, they reported different conclusions. It would be useful to clarify these points.

      In any case, the d-prime of the identifying dimension appears to be a key parameter. It would be great to constrain this parameter using the empirical data. When the d-prime of the identifying parameter is small, the observer would easily confuse the probed stimulus with the other stimulus in a given trial. This should lead to poor task performance. Thus, it may be possible to directly estimate the value of the d-prime of the identifying dimension based on the observer's performance, and then use this parameter to generate model predictions accordingly.

      (1e) The current model assumes that a large number of samples are generated per stimulus and the brain can manipulate this information to perform the demixing task. It was well documented that visual working memory has a capacity limit (i.e., it can only hold information about a few items); this discrepancy needs to be clarified or addressed.

      (2) How well the computational model can explain the experimental data remains not entirely clear

      The authors show that there exists a parameter regime that can qualitatively explain the experimental finding. They also show that it is possible to fit the model to the data to explain the bias patterns. However, given that the model is flexible, it would be stronger if the authors could show that the same parameters that explain the biases could also explain other aspects of the behavior, for example, the magnitude of the errors.

      In other words, the model is not well constrained in the way it was tested in the paper. But it should be possible to improve it. First, if the noise parameter in the model is determined by the stimulus variability, one can determine it directly based on the external noise in the stimuli (discussed also in 1b) and see what prediction it leads to. Second, from the behavioral data, it may be possible to estimate the noise for the identifying dimension. Doing so will help better constrain the model.

      It would also help if the authors could report the best-fitted parameters from the experimental data. From these parameters, one can simulate synthetic data and apply the demixing model to see if the error distribution of the simulated observers is indeed similar to the experimentally measured error distribution. That way, one can check whether the fitted parameter explains the observer's behavioral performance beyond the biases.

      Other comments:

      (1) How does the model account for the swap errors? I am not sure I understood the way how the swap errors were treated in the paper. To me, substantial swap errors seem to be a consequence of having low d-prime values for the identifying dimension; that is, if there is only little information to discriminate the identity of the two stimuli, swap errors would be large. However, this possibility didn't seem to be mentioned in the paper.

      (2) Since the solution of the demixing problem was obtained using a numerical procedure based on EM. It would be useful to check whether the initialization has affected the biases obtained.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the origins of inter-item biases in visual working memory. The authors proposed a computational model where overlapping memory signals are disentangled, inducing memory biases that depend on relative noise levels across items. The key theoretical advance is the prediction that bias direction depends not only on absolute memory noise but on the relative noise levels of target and non-target representations. Using four experiments with color mosaics whose color variability manipulates memory precision, the authors report that biases reverse as a function of relative noise in a manner predicted by the model.

      Strengths:

      The manuscript is clearly written and theoretically motivated. The experiments are well designed and provide converging evidence for a distinctive and non-intuitive prediction of the proposed model. I found the central result compelling: independently manipulating target and non-target noise leads to qualitatively different bias patterns, consistent with the model's prediction that relative noise is a key determinant of bias direction.

      Weaknesses:

      The main limitation is that the evidence establishes consistency of the data with the proposed Demixing Model, but does not demonstrate that the model provides a unique explanation of the data. Although the manuscript argues that dominant theories struggle to account for the observed reversals, no formal comparison with alternative computational frameworks is presented. In addition, model fitting results are reported only briefly, making it difficult to evaluate fit quality at the level of individual observers.

    4. Author response:

      Reviewer #1 (Public review):

      Strengths:

      Overall, the computational model tested in the paper is novel and interesting.

      The demixing framework represents an appealing hypothesis that deserves further investigation.

      The current paper provides new empirical data showing that the target stimuli with the same absolute noise level can be either repelled from or attracted to non-target items, depending on the relative noise levels. The observation that biases depend on the relative noise levels is by itself an interesting one, and is consistent with the prediction of the demixing model.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      While this manuscript contains interesting new experimental observations and theoretical ideas, it has several substantial problems in its current form, which limit the conclusions that can be drawn. The description of the computational model is too brief. The key modeling assumptions need to be better motivated and explained. As the computational models generate different predictions in different regimes, it is a bit difficult to evaluate how well the experimental data support the model at a more quantitative level. Also, the results focused on studying the biases in the behavior; it is unclear whether the model can fully explain the behavior data (such as error distributions or behavioral precision).

      We agree that the model description should be expanded and that quantitative agreement with the data should be assessed more thoroughly, and we plan to address this during the revision. In the initial version of the manuscript, we aimed to highlight the qualitative agreement of the data with the novel and counterintuitive predictions by the model. While the reviewer is correct that the model "generates different predictions in different regimes," the particular predictions we test (the interaction between the target noise level and the parity of the target and non-target noise levels in Experiments 1-3, and the effects of non-target noise when the target noise is held constant in Experiment 4) hold across regimes (Figure S1 shows this for the former prediction). We aim to further expand on this point in the revision.

      Major concerns:

      (1) Concerns/suggestions regarding the computational modeling

      The current paper seeks to test the predictions of the demixing-based computational model proposed in reference 22. There are several problems with the modeling component in the current paper.

      (1a) The description of the model is too brief and difficult to understand. Although the model was proposed in reference 22, it would still be beneficial to provide more details of the model so that readers can understand and appreciate the strengths/limitations of the model.

      The generative model and the inference procedure could be better explained to better link the model to the behavior. In particular, how was the observer's behavioral report in each trial modeled? This requires more explanation because currently the demixing procedure estimates four parameters for a given trial, yet for a given trial, only one behavioral report was produced (e.g., current Experiment 1), or two reports were produced sequentially (e.g., current Experiment 2).

      We will provide more details about the model and how it was fitted to the data. Please note that the model parameters were fitted per subject and condition, not per trial: 2 hue noise parameters,  and , corresponding to the noise of the target and non-target item across 4 noise combination conditions, plus a shared identifiability noise, , across conditions, determining the discriminability of the items along the identifying dimension. This strongly limits model flexibility as only 3 parameters (including the shared  across conditions) are used to create the bias curve for each subject in each condition.

      (1b) Key modeling assumptions need better justification.

      One such key assumption is that on a given trial, each stimulus triggers many samples (or approximately, an entire response distribution), rather than a single sample. This assumption deviates substantially from prior work on ideal observer models. It was not clear whether this assumption is realistic. For the type of stimuli used in the current experiments, perhaps one can argue that each pixel corresponds to one sample of brain activity, thus collectively each stimulus should trigger many samples of activity in the brain. If this were to be the case, it would have two implications. First, the noise parameter in the model should be directly related to the magnitude of the stimulus noise. Thus, one should be able to plug these experimentally-controlled parameter values into the model to directly generate predictions about the biases. Second, when using stimuli with no stimulus variability (e.g., simple grating stimuli), the predicted biases should change. However, it wasn't clear whether this would hold experimentally, i.e., using gratings would lead to different biases or no biases.

      If the variability of the samples for a given stimulus involves neural noise, it would be useful to justify why it is reasonable to consider that many samples were generated per stimulus.

      We are grateful to the reviewer for raising this point, and we will provide more details on it in the revision. In brief, we believe that it is the standard ideal observer assumption of one sample per trial that is unrealistic and works only in cases when there is a single signal source, so that the samples can be simplified to a single average. Consider that determining a stimulus value is a similar problem for an ideal observer to the one that a researcher who aims to decode neural data from populations of neurons (or fMRI voxels) has to solve. Different populations of neurons would provide responses that match different stimuli – in essence, creating different samples in an ideal observer framework. Thus, even without external noise, the demixing problem would be present when there is more than one stimulus, but internal noise is much more difficult to control, so in our experiments, we used multi-colored stimuli.

      (1c) As mentioned in (1b), the model assumes that on each trial, a large number of samples was generated. It would be useful to study and report how the prediction would change when the number of samples generated per stimulus is small. In particular, what happens when each stimulus only generates one measurement? This might be useful for interpreting previous experiment results with grating stimuli.

      This is an interesting point that we aim to address in the revision.

      (1d) Reference 22 studies how the predicted biases depend on the d-prime of the identifying dimension and found that the pattern of the biases varies substantially depending on the information available for the identifying dimension. However, the current paper didn't really discuss this important point. It is also unclear what parameters the authors used for the d-prime of the identifying dimension. Was it fitted directly to the data? The Methods section has some description on the "identifiability dimension", but it was a bit obscure.

      Intuitively, when the d-prime of the identifying dimension is very large, the demixing problem becomes irrelevant. In this case, there should not be any biases induced by demixing. In the case of the d-prime for the identifying dimension is 0, the problem should reduce to the simplified 1-d problem studied in reference 22. If my reading of reference 22 was correct, they reported different conclusions. It would be useful to clarify these points.

      We are grateful for the suggestion to expand the discussion of this point and will do so in the revision. The reviewer is correct that for very large d-prime in the identifying dimension, the demixing problem solution is trivial. However, the 2D case does not resolve to the 1D case when d-prime reaches zero. This is because the identifying dimension is still used to identify which item to report—unlike in the 1D case, when the reported dimension is the same as the identifying one. Consider what happens if the observer in our task does not remember at all which stimulus was left and which was right. It would report the other item in 50% of cases, leading to a strong attractive bias.

      In any case, the d-prime of the identifying dimension appears to be a key parameter. It would be great to constrain this parameter using the empirical data. When the d-prime of the identifying parameter is small, the observer would easily confuse the probed stimulus with the other stimulus in a given trial. This should lead to poor task performance. Thus, it may be possible to directly estimate the value of the d-prime of the identifying dimension based on the observer's performance, and then use this parameter to generate model predictions accordingly.

      We apologize for the confusion. We constrain the discriminability of items in the "identifying" dimension using the  parameter that determines the noise in that dimension for both items. The means in this dimension are fixed at an arbitrary value, as means and noise are interchangeable when considering discriminability. We will revise the description of the fitting procedure accordingly. Regarding the use of the same values in predictions, while possible, we prefer to keep predictions separate from fitting to avoid them becoming postdictions. The curves for the fitted model in Figure 2 already illustrate what the model predicts under the fitted parameter values.

      (1e) The current model assumes that a large number of samples are generated per stimulus and the brain can manipulate this information to perform the demixing task. It was well documented that visual working memory has a capacity limit (i.e., it can only hold information about a few items); this discrepancy needs to be clarified or addressed.

      We are grateful to the reviewer for raising this point, which we will address in the revised discussion. Briefly, we believe that the number of samples in the ideal observer model does not correspond directly to the working memory “slots”.

      (2) How well the computational model can explain the experimental data remains not entirely clear

      The authors show that there exists a parameter regime that can qualitatively explain the experimental finding. They also show that it is possible to fit the model to the data to explain the bias patterns. However, given that the model is flexible, it would be stronger if the authors could show that the same parameters that explain the biases could also explain other aspects of the behavior, for example, the magnitude of the errors.

      It would also help if the authors could report the best-fitted parameters from the experimental data. From these parameters, one can simulate synthetic data and apply the demixing model to see if the error distribution of the simulated observers is indeed similar to the experimentally measured error distribution. That way, one can check whether the fitted parameter explains the observer's behavioral performance beyond the biases.

      We are grateful to the reviewer for raising this point. We both agree and disagree with the reviewer here. The predictions reported come from an earlier paper describing the model (ref. 22). In our opinion, this represents a pure hypothesis-driven approach, where a prediction is formulated first and then tested with subsequently collected data. The model we test is normative, not descriptive; its goal is not to fit the data as closely as possible, but rather to make predictions about internal brain mechanisms. We do not suggest, for example, that demixing is the sole source of biases, so the resulting bias pattern might differ significantly from the predictions. That the model fits the data is, therefore, an additional bonus. At the same time, we agree that it is interesting to test whether the model can explain other parameters of the data. Note that our current fitting procedure was not geared toward this; we optimized the model to explain only the bias curve. In the revision, we aim to test whether the model can also explain the error variability.

      In other words, the model is not well constrained in the way it was tested in the paper. But it should be possible to improve it. First, if the noise parameter in the model is determined by the stimulus variability, one can determine it directly based on the external noise in the stimuli (discussed also in 1b) and see what prediction it leads to. Second, from the behavioral data, it may be possible to estimate the noise for the identifying dimension. Doing so will help better constrain the model.

      External noise accounts for only a portion of the total noise, as evidenced by behavioral errors. Even for a single item, the total noise consists of the amount of information the observer samples from the stimulus, the variability of these samples (external noise), and early (applied to each sample) and late (applied after integration) internal noise. Therefore, external noise alone might not constrain the model in the right regime. Regarding the identifying-dimension noise, as noted above, we do constrain it with the data. However, we aim to explore these points further in the revision.

      Other comments:

      (1) How does the model account for the swap errors? I am not sure I understood the way how the swap errors were treated in the paper. To me, substantial swap errors seem to be a consequence of having low d-prime values for the identifying dimension; that is, if there is only little information to discriminate the identity of the two stimuli, swap errors would be large. However, this possibility didn't seem to be mentioned in the paper.

      We apologize for the confusion. We will further clarify and perhaps reassess the treatment of swap errors in the revision. The model itself produces swap errors when the stimuli sources are misidentified.

      (2) Since the solution of the demixing problem was obtained using a numerical procedure based on EM. It would be useful to check whether the initialization has affected the biases obtained.

      Indeed, this is a valid point, and it's why we use a multi-initialization strategy. For each simulation of a single trial sample set (e.g., 100 random samples), we use a large number of initial points (50 in the initial submitted manuscript) to ensure the obtained EM solution is truly optimal. Additionally, we conduct a large number of simulated trials (10,000 for each parameter combination) to ensure the accuracy of the bias distribution we obtain.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the origins of inter-item biases in visual working memory. The authors proposed a computational model where overlapping memory signals are disentangled, inducing memory biases that depend on relative noise levels across items. The key theoretical advance is the prediction that bias direction depends not only on absolute memory noise but on the relative noise levels of target and non-target representations. Using four experiments with color mosaics whose color variability manipulates memory precision, the authors report that biases reverse as a function of relative noise in a manner predicted by the model.

      Strengths:

      The manuscript is clearly written and theoretically motivated. The experiments are well designed and provide converging evidence for a distinctive and non-intuitive prediction of the proposed model. I found the central result compelling: independently manipulating target and non-target noise leads to qualitatively different bias patterns, consistent with the model's prediction that relative noise is a key determinant of bias direction.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      The main limitation is that the evidence establishes consistency of the data with the proposed Demixing Model, but does not demonstrate that the model provides a unique explanation of the data. Although the manuscript argues that dominant theories struggle to account for the observed reversals, no formal comparison with alternative computational frameworks is presented. In addition, model fitting results are reported only briefly, making it difficult to evaluate fit quality at the level of individual observers.

      We agree and we aim to provide a comparison with alternative models and an expanded description of the fitting results in the revision. Note, however, that the majority of existing models are descriptive, while we believe that as a normative model, the Demixing Model should be compared with other normative models, thus limiting the selection of competitors significantly.

    1. eLife Assessment

      This useful study reports on lifespan extension in C. elegans males that carry a mutation in a gene for an insulin receptor; while the observations are striking, the strength of evidence is currently incomplete. The central claim of male-specificity is undermined by the absence of direct hermaphrodite healthspan comparisons, and the reliance on a single mutant allele leaves open the possibility that background mutations, rather than daf-2 loss-of-function, drive the phenotype. Methodological details critical for reproducibility are also lacking, particularly regarding male housing density, censoring of plate-leaving animals, and the adequacy of replication for the key epistasis experiment. The work could be substantially strengthened by a targeted set of additional experiments and fuller engagement with the existing literature on sex-specific aging in C. elegans.

    2. Reviewer #1 (Public review):

      Summary:

      This paper provides interesting observations about the effects of a classical mutation in the daf-2 insulin-like receptor in male C. elegans. The observations are a contribution in and of themselves; however, the conclusions reached about these observations are not supported by the work presented. Most importantly, male-specific effects on healthspan measures are asserted without direct comparison to hermaphrodites. Perhaps more fundamentally, essential features of the methods and experimental design are lacking, which makes formal assessment of the results impossible, especially given our knowledge of negative male-male interactions, which have gone completely unacknowledged here. Indeed, there is a general lack of context for known sex differences in C. elegans, especially in terms of the core elements of longevity, which are presented here as entirely novel but in fact are not.

      Major comments:

      (1) The main overall criticism of the premise of the paper is that it lacks a clear hypothesis that would lead to explicit experimental tests. Instead, many of the results are observational, and the conclusions reached go beyond the actual experiments conducted. The goal should be explicit and consistent between the introduction/ discussion, and the findings should directly address the goal.

      The overall focus appears to be that daf-2 males have an extended lifespan for reasons that are different from hermaphrodites. This conclusion is apparently based on the observation of lipid reserves in mutant animals. However, none of the healthspan measures were conducted in parallel with identical measures in hermaphrodites. How can the authors then claim that males are unique? This is especially problematic since other studies have demonstrated that daf-2 hermaphrodites also have altered lipid composition (Vrablik 2015 Biochim Biophys Acta; Horikawa 2010 Mol Cell Endocrin).

      (2) The authors make unwarranted claims about causation from observational data that is correlative in nature. Again, they claim that male longevity is caused by increased lipid reserves. This may in fact be the case, but there is no evidence to show that this is causal, only that lipid reserves are increased in mutant animals. Causation requires an actual experiment, in this case, disrupting lipid maintenance in daf-2 males (e.g., Lapierre 2013 Autophagy). Their conclusions are consistent with their results, but their conclusions are much too strong given the nature of the evidence, especially given the concerns about proper comparisons to hermaphrodites.

      (3) With these concerns in mind, all conclusions related to male-specific effects should be statistically tested using a sex-by-treatment interaction term in the statistical model. This is obviously impossible for the healthspan data, but for lifespan, this can be directly tested using (genotype x sex interaction in the CPH analysis). Further, it is unclear why each of the replicates is shown separately in Figure 1.

      It is nice that the authors do not directly pool them, as most longevity studies do, but the replicate effects can be included in a more comprehensive model, which would yield an appropriate "average" effect curve.

      (4) There is an inadequate review of pre-existing literature and findings that predate the observations presented here. While this is not an issue in general, the authors present their work as entirely novel when it is not.

      In addition to Gems and Riddle (2000), which is tangentially cited in the discussion, the following papers should be cited and discussed in the introduction to clarify what is currently known and what remains to be explored:

      Partridge and Gems (2002) Mechanisms of aging: public or private?

      McCulloch and Gems (2007) Sex‐specific Effects of the DAF‐12 Steroid Receptor on Aging in Caenorhabditis elegans

      Hotzi et al (2018) Sex‐specific regulation of aging in Caenorhabditis elegans

      Al-Saadi et al (2025) Disruption of the insulin signaling pathway in C. elegans dramatically increases male longevity and enhances reproductive health late in life

      In addition, the authors assert that the study of sex differences is unstudied. If the authors are specifically referring to the sex differences in aging research, they should explicitly state that and revise their language to reflect that it is "understudied" rather than "unstudied". But as stated below, there are many studies that look at sex-specific differences in behavior, physiology, development, etc. This is most important in the context of sexual conflict, of which there are many studies that are directly relevant to the work presented here. The authors are encouraged to review some of these papers.

      (5) This is particularly important in the context of how the experiments presented here were actually conducted. The methods are inadequate to assess this, and the results would therefore be impossible to replicate in the absence of additional details. Exactly how many individuals were raised on each plate during the longevity assays (and other work) is critical to understanding the results of this study. This is because males have direct, chemically and physically mediated negative impacts on one another (see many papers from the Brunet and Murphy labs). Further, it is not even clear whether males and hermaphrodites were reared separately from one another. Males are known to leave plates without hermaphrodites, which requires appropriate inclusion of censoring criteria in studies such as these. It is unclear whether and how this was handled. Censoring is an essential feature of any longevity study and so needs to be explicitly described in the statistical methods.

      The methods describe the use of heat shock to induce the production of males, but it is unclear which generation is being used here. Ordinarily, males would be induced, and then male populations would be maintained by forced mating (picking to ensure that there is a high relative frequency of males) for several generations to eliminate any carryover effects of the heatshock itself. Were the heatshock males put directly into the longevity assays? If so, were hermaphrodites subject to identical treatment? This is confusing, and a potentially critical confound is not performed correctly.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript presents interesting observations regarding the exceptional longevity and improved healthspan of male daf-2 mutants. Given the comparatively limited focus on male aging in C. elegans, the study provides a potentially useful characterization of sex-specific effects associated with reduced IIS signaling.

      Strengths:

      The 4-fold increase in lifespan of male daf-2 mutants is a striking and unexpected observation. The altered fat metabolism between older daf-2 mutant males and hermaphrodites provides further evidence of sex-specific effects.

      Weaknesses:

      (1) A major limitation of the current study is that the conclusions rely primarily on a single daf-2 allele. It would strengthen the manuscript to validate at least the major observations using an independent daf-2 allele or through daf-2 RNAi. This is particularly relevant for the proposed male-specific enhancement of longevity and healthspan, as it remains unclear whether the observed effects broadly reflect reduced IIS signaling or may be influenced by allele-specific effects or background mutations.

      (2) The methods for male lifespan assays require additional detail. Although the authors state that males were generated and transferred every three days, it is not clear whether males were maintained singly or in groups, how many males were placed per plate, or how many were censored by fleeing. These details are particularly important for male aging assays, as male lifespan in C. elegans is known to be influenced by social interactions. Factors such as population density can affect survival and healthspan measurements. Clarifying these procedures would improve reproducibility and interpretation of the reported male-specific lifespan effects.

      (3) Because reduced IIS signaling in daf-2 mutants is known to alter metabolism, physiology, and potentially male-derived signaling, it would be interesting to determine whether the enhanced longevity of daf-2 males is influenced by altered male-male interactions or resistance to male-associated toxicity. In this context, clarification of whether lifespan assays were performed with grouped or individually maintained males would be valuable. If not already tested, lifespan analysis under isolated single-male conditions could help distinguish intrinsic longevity effects from potential contributions of population-dependent signaling or social interactions.

      (4) In Figure 2D, the body-length measurements in WT males appear somewhat unexpected, particularly the apparent increase between Day 14 and Day 20. Since adult worms are not typically expected to exhibit substantial growth at advanced ages, additional clarification regarding the measurement methodology would be helpful, including confirmation that the scale bars and image scaling were applied consistently across conditions.

      (5) The use of palmitic acid barriers following Beydoun et al. (2024) is appropriate; however, it would be helpful to clarify whether WT and daf-2 males exhibited comparable fleeing behavior under these assay conditions. Because male worms are highly prone to plate leaving and censoring, genotype-dependent differences in fleeing behavior could potentially influence survival analyses and the number of censored animals. In addition, as Beydoun et al. primarily characterized these barrier conditions using hermaphrodites, it would be useful to clarify whether comparable barrier effectiveness was observed in male lifespan assays.

      (6) Separately, Beydoun et al. (2024) also reported that palmitic acid barrier conditions can influence body-size measurements, whereas PEG-based barriers did not show similar effects on body size. It would therefore be useful to know whether comparable body-length trends were observed under alternative barrier conditions, particularly given the unexpected increase in WT male body length at later ages.

    4. Reviewer #3 (Public review):

      This manuscript reports a striking sex-specific effect of the daf-2(e1370) mutation on C. elegans lifespan. The authors show that male daf-2 mutants exhibit dramatically extended lifespan relative to wild-type males, wild-type hermaphrodites, and daf-2 hermaphrodites. The study also demonstrates increased lipid accumulation in these long-lived males, which is increased further over time, improved late-life motility, enhanced oxidative stress resistance, and a requirement for the downstream effector of daf-2, daf-16, for the longevity phenotype.

      The interest of the work is the magnitude and consistency of the lifespan effect. The authors report large increases in both median and mean lifespan in daf-2 males across independently replicated experiments. This is further supported by healthspan analyses and their finding that male daf-2 mutants maintain improved motility and stress resistance, which argues against the interpretation that lifespan extension merely reflects prolonged frailty. The genetic epistasis experiment demonstrating loss of the longevity phenotype in daf-2;daf-16 double mutants provides evidence that the effect depends on canonical insulin/IGF-1 signalling.

      The main limitation is that at least the first figure is rather an incremental increase on previous work examining the lifespan of daf-2 males, although the authors do indeed show that the effects can be much larger (or more 'plastic') than those previously published. While these findings are potentially important, the manuscript would certainly benefit from a more extensive discussion of how the results compare with prior studies of daf-2 mutants and male longevity, including possible explanations for the apparent discrepancies.

      The epistasis experiment shows that this exceptional longevity requires the expression of daf-16. However, in contrast to the initial experiments (Figure 1) that show three replicates of the lifespan experiment (the standard in lifespan work in this model), it appears that the daf-2;daf-16 experiment has only been performed once.

      In addition, the lifespan data for hermaphrodite daf-2 mutants appear somewhat unusual. Although the mean lifespan is increased, the median lifespan is reported to be only modestly greater than that of wild-type hermaphrodites. I know that this mutant can give lifespan curves that look like this, but either the use of another allele or of the experimental conditions and how these values compare with previously published daf-2(e1370) datasets would help readers interpret the magnitude of the male-specific effect.

      The lipid phenotype is intriguing. It would be interesting to expand this to examine somatic vs embryonic fat. In addition, I noted that in the methods section, the authors use palmitic acid to stop the male worms 'fleeing' the plates; is it possible to rule out the possibility that the daf-2 mutants are simply eating and metabolising/storing this fatty acid barrier differently than their wild-type counterparts? This would be worth considering and controlling for, particularly as male C. elegans have been shown to have dramatically altered metabolic transcriptional profiles. If indeed this increased lipid is responsible for the extreme longevity of the daf-2 mutant males, it would be desirable to try to link this mechanistically to the phenotype.

      Overall, the evidence convincingly supports the conclusion that male daf-2(e1370) mutants are exceptionally long-lived under the conditions tested and that this phenotype requires DAF-16. The work has the potential to make an important contribution to understanding sex-specific regulation of ageing, although further contextualisation within the existing literature would strengthen the manuscript.

    1. eLife Assessment

      This valuable study provides novel evidence that congenital aphantasia is associated with structural differences in frontotemporal and cingulate systems, with relative sparing of early visual regions and major visual pathways. The multimodal structural imaging approach is carefully implemented and will be of interest to researchers studying mental imagery and aphantasia. However, the strength of evidence is incomplete because the data cannot adjudicate between alternative cognitive interpretations, and the multiple discovery streams make the findings better viewed as key exploratory evidence, rather than as establishing a definitive structural phenotype of aphantasia.

    2. Reviewer #1 (Public review):

      Summary

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (3) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

    3. Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths:

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (4) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

      We will add two recent in-press Cortex papers to the Discussion. One provides lesion-based double-dissociation evidence against V1 as a necessary causal substrate of visual imagery. The other shows that aphantasic individuals can display visualizer-like oculomotor patterns during mental map exploration despite reporting little or no imagery vividness. Together, these studies help clarify our interpretation of our null V1 findings and structural effects in higher-order brain regions, which are consistent with aphantasia involving altered integration or access rather than a primary V1-dependent imagery deficit.

      Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

      We thank the reviewer for this thoughtful and constructive comment. Lack of introspective report of voluntary imagery is arguably the defining signature of aphantasia. This motivated us to primarily interpret our anatomical findings in a broader cognitive context of higher-order control, internal monitoring, and conscious access in aphantasia. We expect that a reliable behavioural test measuring imagery sensitivity and accessibility would allow us to direct link these findings to individual imagery ability. Nevertheless, to our best knowledge, this kind of test on imagery is still missing. Instead, our findings point to some plausible structural signature or brain regions that may be related to conscious imagery, which motivate future studies to examine their direct or causal roles. We agree with the reviewer, future studies should test the relationship between these anatomical structures and the accessibility to internal representation, together with related aspects of internal monitoring. We will therefore add a dedicated paragraph to discuss the plausible cognitive mechanisms during the revision.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

      We are grateful to the Reviewer for the constructive and thoughtful assessment of our manuscript. In response to the reviewer’s comments, we will revise the manuscript to clarify the dependency between diffusion-derived analysis streams, to state more explicitly the biological limits of diffusion MRI metrics, to add a null-network sensitivity analysis for the clustering coefficient findings, to include confidence intervals for reported effect sizes, and to temper the interpretation of the cortical thickness result. We will also revise the Abstract and Discussion to better reflect the exploratory nature of the study and to frame the findings as hypotheses requiring replication in larger independent samples. We believe that these revisions will make the manuscript more balanced, transparent, and appropriately cautious, while preserving the central conclusion that congenital aphantasia is associated with structural differences centered on higher-order frontotemporal and cingulate systems.

    1. eLife Assessment

      This valuable study identifies Gcn5 as a regulator of blood cell development in the Drosophila lymph gland, with links to autophagy and nutrient-sensing mTORC1 signalling. The evidence is solid that altering Gcn5, autophagy genes and mTORC1 activity perturbs blood cell homeostasis, and the revised manuscript adds helpful genetic and quantitative analyses. However, the evidence for a clean linear Gcn5-mTORC1-TFEB/autophagy pathway is insufficient, because several cell-type-specific phenotypes remain difficult to reconcile and the pathway logic relies on different genetic tools, cell populations and pharmacological perturbations.

    2. Reviewer #1 (Public review):

      In their manuscript Arjun et al. investigate the role of the histone acetyl transferase Gcn5 in controlling drosophila blood cell homeostasis in the larval lymph gland. Using gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland cell populations, they show that these manipulations impact (but in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation, DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and show that reducing the expression of the autophagy machinery affect blood cell differentiation. By using drugs as well as genetic approaches to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR, but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      Overall, the main conclusions are sound but interpreting several lines of experiments and results remain complicated. Consequently, the overall picture of the role of Gcn5 in Drosophila larval lymph gland development, and its relationship to mTOR and autophagy, remains unclear.

    3. Reviewer #2 (Public review):

      Summary:

      Drosophila haematopoiesis has been shown to be governed by a number of signalling pathways such as JAK/STAT and Dpp. This important study shows a role for nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it TFEB activity thus acting as a fine-tuning mechanism which maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy has a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weakness:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections do not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. And the modulations of Gcn5 in PSC also has variable effects across hemocyte subpopulations which is not explored in the manuscript. Interestingly, also the domain deletion constructs show differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained. While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4. The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      Comments on the revised manuscript:

      Overall, the revisions make the narrative more coherent. The authors have also added data which substantiates their conclusions.

      However, in some instances, the authors are not clearly able to explain the discrepancies in the data (Gen-5 depletions under the Hml-Gal4 in the whole larval lysates remove p62 completely) which is not ideal.

      A query regarding the discrepancies in the immunofluorescence data: The authors have removed the IF data which suggested that there could be differences in the shuttling of Gcn5 between the nucleus and cytoplasm. The authors suggest that immunofluorescence issues are at the root of these variable results, but the reviewer wonders whether there could be further unexplored mechanisms re: shuttling that is unexplored here and would have been potentially novel.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In their manuscript, Arjun et al. investigate the role of the histone acetyltransferase Gcn5 in the control of drosophila blood cell homeostasis in the larval lymph gland. They use gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland populations to show that these modulations impact (in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation or DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and they show that decreasing the expression of the autophagy machinery increases blood cell differentiation. Using drugs to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      While the authors did a lot of experiments and good quantifications of the blood cell phenotypes, many results do not make much sense or do not bring valuable information about Gcn5 mode of action. Several conclusions of the manuscripts are not backed by solid data (e.g. that Gcn5 action is mediated by TFEB and the autophagy machinery) and different aspects of the literature are not well taken into consideration. Some results (such as the validation of the knockdown and overexpression of Gcn5) seem flawed. There are some concerns about the results obtained with gcn5 zygotic mutants and an interpretation of the phenotypes observed upon manipulation of Gcn5 expression in different cell types is missing.

      We have now performed several experiments to address the comments raised by the reviewer and have also provided possible explanation of the phenotypes in cases where it was lacking.

      Important revisions are needed to improve the quality of the manuscript and confirm the authors' findings.

      Reviewer #2 (Public Review):

      Summary:

      Drosophila hematopoiesis has been shown to be governed by a number of signaling pathways such as JAK/STAT and Dpp. This important study shows the role of nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it, TFEB activity thus acting as a fine-tuning mechanism that maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy have a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub-sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weaknesses:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections does not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. The modulations of Gcn5 in PSC also have variable effects across hemocyte subpopulations which is not explored in the manuscript.

      We have now explained and discuss why Gcn5 modulation could be affecting the PSC size. Please check Discussion section Paragraph 1 line 10 onwards.

      Interestingly, also the domain deletion constructs show a differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained.

      Currently, with our observations all that we can comment about that data is that expression of domain deletion mutants causes aberrant hematopoiesis indicating a dominant negative phenotype since they are expressed in the wild type genetic background. Beyond this, we will be exploring mechanistically how these domains are functioning during hematopoiesis in future studies. We have already described the dominant negative effect in the text: Discussion Section Paragraph 3.

      While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4.

      We have now revised Figure 1 and have only included the images with Collier-Gal4 and Dome-Gal4 which clearly shows both the niche cells, Dome-positive progenitors and Dome-negative cells of the primary LG lobe essentially showing that Gcn5 is expressed throughout the primary LG lobe. In Fig. 1C-F’, Gcn5 expression is both in the nucleus and cytoplasm as this molecule shuttles between cytoplasm and nucleus. The staining pattern with the other Gal4 could be due to problems in the immunofluorescence protocol and acquisition parameters. We have now removed those images from Figure 1. Please check revised Figure 1.

      The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      We have used a pan-hemocyte promoter for the autophagy analysis to investigate if Gcn5 regulation over autophagy is a hemocyte specific effect which we indeed see. We have removed the western blot data now in the revised manuscript where we looked at Atg8 and p62 levels in whole larval lysates when Gcn5 was perturbed using hemocyte driver as the results were puzzling and difficult to comprehend given the complete absence of a p62 band in Gcn5 knockdown conditions. Also, it’s worth noting that Hml-Gal4 is also active in the LG hemocytes. We did 2 alternate promoters for prohemocytes to cross-validate some of our results and the chemical modulators experiment was done since effects like mTOR inhibition/nutrient sensing effects are systemic and hence such modalities were employed.

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      It is a possibility that Gcn5 perturbation could be affecting lymph gland size although we have not seen any consistent trend that would point towards this phenotype either upon knockdown or over-expression. We believe Gcn5 controls blood cell differentiation phenotypes strongly via mTORC1. But in order to answer reviewer’s comment we have now re-analyzed our crystal cell differentiation data particularly and quantitated it and represented it as crystal cell differentiation index for dome-Gal4 specific Gcn5 modulation and for the data with genetic modulation of mTORC1 pathway. Please see Fig 3P and S10J for the revised analysis.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      We thank the reviewer for their useful critique. We have now addressed this concern and we have genetically perturbed the mTORC1 pathway in the progenitors using both abrogation of TORC1 via depletion of Tor or Raptor or by activation using over-expression of Rheb. We have now included this data as Supplementary figures – Fig S10 and S11 and have described it in the results section. Please see results section “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The abstract could clearly be improved. It does not make a clear presentation of what is new in the manuscript. The conclusions that Gcn5 function in the lymph gland is mediated by the autophagy machinery and the acetylation of its non-histone target TFEB are not grounded and purely circumstantial. The implication of mTOR and nutrition in drosophila larval blood cell homeostasis has already been studied but not mentioned here. Most of the time the authors do not provide any possible explanation about the phenotypes they observe and how they fit with the current literature. Several pieces of results are of serious concern.

      We would like to thank the reviewer for their feedback. We have revised the abstract and have incorporated the insights obtained from our study. We have now included relevant literature that talks about the implication of mTOR and nutrition in Drosophila larval blood cell homeostasis (Please see Introduction section Paragraph 2 in the manuscript). We have also noted the input of the reviewer on many phenotypes lacking any description of a possible explanation. We have worked on results section to provide possible explanation and speculation wherever relevant.

      In the introduction, the authors do not provide an up-to-date and accurate presentation of the field. For example, they could use much more recent and comprehensive reviews since Evans et al. 2003. (eg. MID: 30733377 or 35887113). Their choice for signaling pathways involved in Drosophila blood cell progenitors seems very much biased for lead author self-citation rather than more directly related citations. It is surprising too that the authors failed to mention a series of publications on Akt/mTOR and nutrient sensing impact on drosophila larval blood cells (PMID: 22951642 ; 22911822 ; 22407365 ; 22510984). Along the same line, there are already several reports on autophagy genes implicated in Drosophila hematopoiesis and blood cell functions (PMID: 23406899; : 33560224 ; 20498061 ; 37623416). The introduction on GCN5 is a bit of a catalogue and should be streamlined- citing a recent review would be useful (PMID: 32735945). Again, the authors fail to cite publications showing that Gcn5 levels can be modulated by nutrition (PMID 27022023; 27874008) and they do not mention that amino acids starvation or mTOR inhibition leads to a decrease in GCN5 activity / TFEB acetylation (ref 40). Taking into account all the missing information, the novelty of the present manuscript is strongly decreased.

      We would like to thank the reviewer for the detailed suggestions on including the relevant literature that are appropriate and relevant to be mentioned in the context of the observations in our manuscript. We have now included these references and have cited them as per reviewer’s suggestions. Please check Introduction section paragraph 2.

      Results

      While there is little doubt that Gcn5 is expressed in the entire primary lobes based on Fig 1C-F, the quality of the staining in G-J (especially H, J) is really poor and essentially looks like non-specific background with no clear signal in the nuclei. Better images should be presented. The conclusion of the paragraph ("all cellular populations of the LG") and title of Fig.1 are not fully accurate as the authors do not provide evidence that Gcn5 is also expressed in posterior lobes.

      As per reviewer’s suggestions, since the images in Fig 1A-D’ clearly show that Gcn5 is expressed in the entire primary LG lobe in PSC cells, MZ and CZ; we have removed panels E-H’ which lacked clear nuclear signal. Fig1A-D’ clearly show the nuclear staining pattern of Gcn5. We have also modified the conclusion of the paragraph to say that Gcn5 is expressed in cellular populations of the primary lymph gland lobe accordingly.

      Concerning, Fig S1 and Fig 2, while the analysis seems technically sound, the results are puzzling. The lack of P1 differentiation in gcn5 null heterozygotes is very surprising. The authors should check that this stock does not carry a mutation in nimC1 (for details see: PMID: 23899817) and use other plasmatocyte differentiation markers to confirm their observation (also with the different allelic combinations). I'm also concerned by the levels of plasmatocyte differentiation and crystal cell number in the control line (notably in S1H), which seem very low (and quite variable for P1 as there is a notable difference between S1H and Fig 2H). Moreover, the analysis of the allelic combinations gives rather incoherent results: PCSC cell numbers are affected only in null/hypomorph, whereas differentiation (NimC1 and Hnt), as well as DNA damage, was only increased in hypomorph homozygotes. The authors propose no hypothesis to explain these observations.

      We have now repeated these experiments with the E333st null allele by placing it on a different balancer and we observe homozygotes that are alive till late third instar/early pupal stage as shown before by Carre et al., 2005. We have now included these revised results on the plasmatocyte differentiation status of the E333st heterozygotes and homozygotes (See Fig 2 and Fig S1). We do find P1 positive cells in the E333St heterozygotes unlike earlier. Plasmatocyte and crystal cell numbers in the control line always shows some level of heterogeneity. We have included the wild type control individually with those respective mutants during the experiment hence drawing a cross comparison across two different experiments would not be appropriate. We have now explained the observations obtained on PSC cell numbers (Discussion section paragraph 1). Experiments to check all hematopoietic aspects of the gcn5 null have been done after changing the balancer line and the null mutants overall show a decrease in PSC size and a widespread increase in hemocyte differentiation which could be due to a systemic effect due to various signalling pathways being affected which needs to be investigated and is beyond the scope of this study. This has also been discussed in the Discussion section Paragraph 1.

      Although a side-by-side comparison would have been better suited, it seems that the homozygotes or trans-heterozygotes do not have stronger phenotypes than the heterozygotes as far as crystal cell and DNA damage are concerned, which is rather unexpected. Besides the authors should introduce why they look at DNA damage.

      We agree with the reviewer that for the crystal cell and DNA damage phenotype the homozygotes or trans-heterozygotes do not have a stronger phenotype as compared to the heterozygotes alone but since these are whole animal mutants there could activation/inactivation of various signalling pathways and systemic effects that would be difficult to account for and comprehend here which needs to be investigated further. The only conclusion that we draw from these observations is that Gcn5 is required for maintaining blood cell homeostasis. Regarding DNA damage, we have now included the rationale and supporting literature for why we have studied DNA damage in the context of Gcn5. Please see result section 2 paragraph 1.

      Importantly too, the authors failed to obtain gcn5 E333st/E333st (null) larvae, whereas Carre et al. originally reported that E333st/E333st individuals are viable until the late third instar larvae. I suspect that the stock they use carries additional mutations that need to be eliminated by back-crossing it to control flies for several generations. Of note too, a recent report showed that a deletion of gcn5 (generated by CRISPR) does not prevent adult emergence, challenging the conclusion that gcn5 expression is absolutely required for fly development (PMID: 37545086).

      The reviewer is right in pointing out that E333st homozygotes survive until late third instar as reported by Carre et al.,2005. We have procured the null allele again and used another balancer to obtain homozygotes and we were able to get homozygotes that survived till late third instar as reported earlier. We have now included new data from these homozygotes for all hematopoietic aspects and heterozygotes particularly for plasmatocyte differentiation Please see Fig 2 and Fig S1 and corresponding results section 2 of the manuscript.

      Concerning the validation of Gcn5 knock-down and overexpression: the results are highly dubious. In Fig S2B (hml>Gcn5 RNAi), there is virtually no Gcn5 signal in the primary lobes but hml is normally expressed only in the cortical zone. How is it possible? Similarly, the western blot (which is really too much cropped around the bands of interest) does not show any signal in the hml>Gcn5 RNAi lane (not even some background. According to the Methods section, the western was performed on whole larvae extracts; hml-mediated knock-down can not wipe out its expression in all the tissues. As for the overexpression, flag immunostaining in hml>Gcn5-flag is mostly cytoplasmic (S2E), which doesn't make sense and does not fit with S2C (Gcn5 immunostaining).

      Hml-Gal4 is a pan hemocyte driver and its expression is not limited to the CZ of the primary lymph gland lobe (Banerjee et al., 2019) and recent single cell sequencing data corroborate this that Hml domain is not limited to the cortical zone (Yarikipati and Bergmann, 2026). GFP driven by Hml-Gal4 is spread out across the primary LG lobe which could explain the phenotype of no Gcn5 signal obtained in the immunofluorescence experiment. Regarding the western blotting experiment which was performed on whole larval extracts, we were also puzzled by lack of Gcn5 bands in these lysates upon depleting Gcn5 using Hml-Gal4. We need to systematically probe further to understand expression of Gcn5 in other tissues and organs. We have now removed the western blot data as the data obtained cannot be comprehended at the moment. Regarding the FLAG staining experiment – the staining gave us a cytoplasmic pattern and since Gcn5 is known to shuttle between the cytoplasm and nucleus it is possible that the anti-FLAG staining detected the Gcn5 localizing in the cytoplasm. It is difficult to draw a direct comparison here between the images S2C and S2E as both are different antibodies.

      The initial analysis of Gcn5 level modulation in the prohemocytes, PSC or Hml+ cells is mainly descriptive and the authors do not elaborate on possible explanations based on the current literature.

      We have added a possible explanation wherever required for these respective results on Gcn5 modulation in prohemocytes, PSC and Hml positive hemocytes. Please see result section 3 where we elaborate on possible explanation for the phenotypes observed.

      The structure/function analysis of Gcn5 is based on overexpression of truncated mutants in the prohemocytes using the tep4-GAL4 driver and monitoring PSC cell, prohemocyte maintenance, plasmatocyte and crystal cell differentiation as well as DNA damage. As the overexpression of the full-length protein was made with a different driver (Dome), it is difficult to interpret the data. Nevertheless, no clear message emerges from this analysis and the authors do not reach any conclusion. Thus, the interest of these experiments remains limited.

      The structure-function analysis was largely done to understand which of the domains of Gcn5 upon over-expression results in a dominant negative like phenotype and our analysis shows that expression of some of these domain mutants results in a dominant negative phenotype in the wild type genetic background which we have now stressed upon in the text. However, further mechanistic understanding and in-depth analysis of each of these domains of Gcn5 warrants further separate investigation and is beyond the scope of this study. Please see the end of result section 4 for conclusion and possible explanation.

      The authors then analyze autophagy markers (in hml>Gcn5 LOF or GOF). Contrary to their say, hml-GAL4 is not a pan-hemocyte marker. It would have been interesting to ensure that the effects observed on Atg8 and Ref(2)P in the lymph gland are cell-autonomous- as expected for a direct role of Gcn5 on this pathway. Again, it is very surprising that p62 is not detected in the western blot on whole larval extracts when Gcn5 is knocked down in Hml+ cells only (Fig 5D). Moreover, quantifications on multiple samples will be needed to validate the increase/decrease of p62 and Atg8 as detected by western blot. As for the RT-qPCR (Fig S5), according to the Methods sections, they were made on adult blood cells but this is not explicit in the result section.

      We have corrected the text and mentioned Hml-Gal4 as a hemocyte specific Gal4 shown earlier as Gal4 marking both embryonic and larval hemocyte population (Goto et al., 2003, Yarikipati and Bergmann, 2026). Regarding the Atg8 and Ref (2)P blots – yes, it is surprising to us too that the p62 is not detected in the larval lysates when Gcn5 is depleted using Hml-Gal4. However, this result was consistent over the replicates performed and needs to be further studied. Since this phenotype of complete absence of p62 in larval lysates upon Gcn5 depletion cannot be comprehended and explained, we have removed the western blot data from the figure and have just retained the immunofluorescence data and have also quantified the Atg8 and p62 puncta per cell and included this data in Figure 5, Graphs D and E. For the qRT-PCR we have now included a description in the corresponding results section. Please see result section – result 5 under “Autophagic flux in the Drosophila blood cells is negatively regulated by Gcn5”.

      The knock-down of TFEB or several autophagy genes in the prohemocytes (tep4-GAL4) leads to a rather convincing increase in plasmatocyte and crystal cell differentiation. It would have been interesting though to quantify prohemocyte maintenance, PSC cell number, and DNA damage. Also, the authors should have performed Gcn5 GOF/LOF experiments with the same driver (they present tep>Gcn5 RNAi in Fig 8 but without the proper controls).

      We have now included data for prohemocyte index (Figure S8M) upon knockdown of TFEB and other autophagy genes along with PSC cell number, DNA damage (Supple Fig S8) in the revised manuscript. Please see corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe” for the description of the results.

      The use of chloroquine should be better described. How long was the treatment? Did the authors observe an effect on autophagy in the lymph gland? Chrorloquine also affects lysosomal pH, so it remains to be demonstrated that the effects observed here are only autophagy-related.

      We have now written a detailed protocol for the treatment in the methods section and also mentioned the treatment time which is 16 hours in the results. We have included data to validate the effect of Chloroquine on autophagy by p62 and Atg8 staining in the LG and have quantitated the data (Refer Supple Fig S9) and the corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe”

      Similarly, the use of drugs to activate (3BDO) or inhibit (Rapamycin) mTOR should be better controlled. More generally, given the promiscuous roles of mTOR (and autophagy) in the larvae, tissue-specific manipulations would be better suited.

      We have now perturbed mTOR pathway genetically by activation and in-activation and have studied the effect on blood cell differentiation. Please see Figure S10 and the corresponding result section titled “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” where we discuss the results of genetic perturbation of mTOR pathway.

      Actually, as pointed out above, it has already been shown that modulation of Akt/TOR in hemocytes or amino-acid deprivation affects blood cell homeostasis (see above). The authors should definitely discuss how their results fit with the literature on this subject.

      We have added relevant literature in the introduction section and have also discussed how Gcn5 could fit into this context of nutritional sensing and control of hematopoiesis. Please check revised Introduction section paragraph 2. Also, check discussion section in last paragraph where we have discussed role of Gcn5 in nutrient sensing.

      Again, Gcn5 levels need to be quantified using multiple samples (Fig 7M, N) before concluding.

      Sorry for not including the quantitation earlier but we have now included the quantitation for the blots presented in Fig. 7 M and N.

      Finally, the authors show that 3BDO still induces an increase in blood cell differentiation when gcn5 is knocked-down in tep4+ cells and that Rapamycin still represses differentiation when Gcn5 is overexpressed in Dome+ cells. They conclude that mTORC1 overrides the effect of Gcn5. This seems a far-reaching conclusion given the available evidence.

      We have now toned down the conclusion that we make to accommodate other possibilities which we have been unable to test here currently.

      In particular, in the conditions used, the authors do not necessarily assess the activity/requirement for Gcn5 and mTORC1 in the same cell population.

      Other comments and suggestions:

      The discovery of the SAGA complex is not Grant 1999 but 1997 (PMID: 9224714).

      Ref 30 is not appropriate -nothing to do with HAT.

      GCN5 not only acetylates TFEB but also Atg7 (PMID: 28594263) to limit autophagy.

      Thank you so much for these suggestions. We have made the necessary amendments in the references.

      In the results section, the first paragraph is largely a repetition of the introduction. The same is true for most paragraphs in this section. A shorter (hypothesis-driven) introductory sentence would be more adequate.

      We have now taken the suggestion into consideration and made the necessary change in the results section throughout the manuscript.

      Fig 1: it seems that there is a higher accumulation of Gcn5 in a few cells in the cortical zone. This may correspond to crystal cells and could be easily confirmed.

      We have now checked this aspect. Please see supple fig S5 where we co-stain lozenge-GFP cells containing LG with Gcn5 to check for the accumulation. However, we do not see any accumulation in the Lozenge-positive crystal cells.

      Figure 3: the authors should also quantify the proportion of progenitors (dome>GFP+) in the different conditions.

      We have now done this and added it to the Figure. Please see panel N in Figure 3 and Figure S8M.

      Figure S3: how do the authors explain that Gcn5 knockdown in the PSC reduces plasmatocytes differentiation (but does not affect PSC cell number or crystal cell differentiation)? What could be the origin of the increase in DNA damage (essentially in CZ)? How do they explain that Gcn5 over-expression increases PSC size but does not affect (reduce?) blood cell differentiation?

      These observations need to be investigated further. We currently have no answer to these comments. The signals that are produced by the PSC could be affected due to which we observe these phenotypes like an effect on plasmatocyte differentiation and an increase in DNA damage whereas no effect on PSC cell numbers or crystal cell numbers which needs to be studied further. Also, in the case of Gcn5 over-expression in PSC we do not know how the increased size of PSC controls differentiation. This would need further experimentation and since this paper is not about the role of Gcn5 in PSC exclusively, we will look into this in our future studies. These aspects will be studied in our future follow-up studies as it is beyond the scope of the current manuscript.

      Figure S4: how do the authors explain the non-cell autonomous increase in PSC cell number upon Gcn5 KD/GOF in hml+ cells? How do they explain the increase in crystal cell number in Gcn5 GOF? Is it really cell-autonomous (i.e. all the Hnt+ cells are Hml+?)?

      We have discussed how Gcn5 depletion or over-expression in HmlΔ cells could affect PSC cell numbers. Please see discussion section, paragraph 1. Regarding the crystal cell phenotype - We have now tested if the increase in crystal cell numbers is cell autonomous by driving Gcn5 over-expression using a crystal cell specific driver and we find that the increase is cell-autonomous. Please refer to Supple Fig S5.

      The discussion is lengthy and should be reduced. It does not appropriately consider the current literature.

      We have tried to reduce the length of the discussion and have also added relevant references as per recommendations of the reviewer.

      Reviewer #2 (Recommendations For The Authors):

      (1) In general, it is not clear why in some of the experiments Tep-Gal4 is used to modulate proteins in prohemocytes while in others Dome-Gal4 is used.

      There is no particular reason. These Gal4’s have been used interchangeably as both label the hematopoietic progenitor population. Although recent single cell sequencing data has identified subsets within the progenitors namely core progenitors marked by tep4 largely and dome being a distal progenitor marker (Cho et al.,2020, Girard et al.,2021), in our study perturbations in Gcn5 using either of the Gal4’s results in a similar phenotype.

      (2) Considering alteration in lymph gland size (Figure 2), the number of positive cells should be analysed in relation to total cell numbers or s4ize.

      Although we do not find any visible differences or defects in the overall LG size in various genetic conditions discussed in this manuscript, we have done so for the plasmatocyte differentiation where we have represented it as plasmatocyte differentiation index (relative to the size of primary LG lobe) throughout the manuscript. We have now done this for crystal cell numbers too for critical genotypes in this manuscript and have represented it is as crystal cell index for example please see Figure 2O, 3P, S5G, S10J where these graphs have now been added.

      (3) Figure 1A G-I' does not look like mCD8 GFP expression, but rather cytoplasmic GFP.

      We have made the change in the figure and the corresponding text accordingly.

      (4) One of the main conclusions in the manuscript is that Gcn5 affects autophagy (Figure 5). Here, the puncta need to be quantified (relative to total cell numbers).

      Thank you for the suggestion. We have now quantitated the p62 and Atg8 positive puncta per cell and have represented it as panel D and E in Figure 5.

      (5) Figure 5 D and E show p62 and Atg8 total protein levels in the larvae when Gcn5 is modulated only in the hemocytes. It is surprising that there is a complete reduction in p62 levels across the whole larvae when Hml gal4 is used for the knockdown.

      Yes, we observe a complete absence of p62 in whole larval lysates when Gcn5 is depleted using Hml-Gal4 and we see this across replicates. This result is indeed puzzling to us and difficult to comprehend as to why a hemocyte specific driver would result in such a dramatic change hence we have decided to remove the western blot data as it is difficult to draw a solid conclusion from. We have retained the immunofluorescence data which shows a consistent alteration in autophagy upon Gcn5 perturbation using Hml-Gal4 and we have now included the quantification for the number of p62 and Atg8 positive puncta per cell for the IF data.

      (6) The beta-actin levels in the western blots in Figure 5 are highly oversaturated and do not represent loading control adequately. Also, it looks like there is substantially more total protein in 5D 3rd lane where Gcn5 is overexpressed.

      Thank you for pointing this out. We have loaded equal amount of protein in all the wells so we are unsure why the actin bands look over-saturated. We have now removed the western blot data from this figure as the data is puzzling and difficult to comprehend given a total absence of p62 in whole larval lysates in Gcn5 depletion conditions using Hml-Gal4. Hence, we are just retaining the immunofluorescence data.

    1. eLife Assessment

      This study uses convincing modeling methods and analyses of rich behavioral datasets to investigate the role of attention in value-based decision making; for instance, as when choosing between two snacks. The results are important, as they challenge existing theories that assume that paying attention to an available option biases the eventual choice toward that option. The results suggest that the correlation between attention and decision-making is formed largely after rather than before the (internal) choice process has terminated, a finding that offers an intuitively appealing rethinking of how attention and decision-making processes interact during value-based choices.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the weaknesses raised in the previous round of review.]

      Summary:

      This study examines whether gaze direction actively shapes choice during food preference decisions or whether gaze and choice evolve largely independently until the moment of commitment. The established framework in this context, the aDDM, assumes that gaze causally biases the accumulation of evidence in favour of the fixated item. The authors show convincingly that this model fails to fit key behavioural patterns across several datasets, as do other published models that make the same assumption. The authors propose an alternative model (Post-Decision-Gaze or PDG) in which gaze and decision formation are decoupled: gaze does not influence the decision process, nor is it drawn toward the ultimately chosen item, until after the decision threshold is reached. Only during the motor execution period (after commitment) is gaze directed to the chosen option. They demonstrate that this model fits several observed patterns better than the aDDM and related variants.

      Strengths:

      The work thoroughly considers multiple models and datasets. It advances an interesting alternative perspective on gaze-decision interactions and highlights meaningful shortcomings in existing models. The authors take the time to explain how modelling assumptions produce specific patterns in the data, which is certainly insightful to readers interested in the modelling of value-based decision making.

      Weaknesses:

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

    3. Reviewer #2 (Public review):

      Summary:

      Zylberberg et al. reanalyze eye-tracking and behavioral data to test two predictions of the attentional Drift Diffusion Model, finding that these predictions are not met. Similarly, predictions of normative models (inspired by rational inattention) are not in line with the data, and the authors propose a post-choice model of attention. This model better accounts for the two effects but also does not account for all patterns, so the authors conclude that eye movements most likely reflect both pre- and post-decisional processes.

      Strengths:

      A clear strength is the systematic falsification-based approach of the paper, establishing (partially) new predictions and testing to what extent these are met by extant models and by a newly developed theory. The authors do a good job in providing intuitions behind the effects and the reasons why models such as the aDDM predict them. The paper is of substantial relevance for the field, as it shows that effects pertaining to the last fixation(s) should be interpreted with caution. Another strength is the paper's transparency as the authors clearly acknowledge that their new model does not do a perfect job either.

      Weaknesses:

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

    4. Reviewer #3 (Public review):

      Summary:

      In this study, the authors reanalyzed choice, RT and gaze datasets collected from human subjects performing a food-choice task. They show that models that posit a causal role for attention in shaping the decision-making process fail to account for empirical observations in the data. These include the attentional drift diffusion model (aDDM) and models that derive attention-choice associations from an optimal policy. The authors show that a model that assumes that gazes are directed towards the chosen option after decision commitment captures more (but not all) empirical findings, suggesting that attention may reflect decisions once they are made instead of contributing to their formation. However, this post-decision-gaze (PDG) model failed to capture all aspects of the data, suggesting that gaze may reflect both decisional and post-decisional operations, and existing models are still missing some features of the gaze-directing process. The authors provide convincing evidence that post-decision gaze explains a number of empirical findings in this task.

      Strengths:

      (1) The analyses are generally appropriate, and the conclusions are supported by the data.

      (2) The study was rigorous, as the authors considered a number of alternative possible models for behavior, and evaluated their performance based on a wide range of qualitative predictions (as opposed to exclusively relying on model comparison).

      (3) The proposal that gaze may largely reflect post-decisional processes is interesting, and as far as I am aware, novel.

      Weaknesses:

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

      In particular, the "saccadic execution" parameter appears far too long and too variable to reflect merely execution; instead, it likely includes decisional components. This would make more sense since manual and saccadic planning essentially rely on distinct brain areas, hence it seems unrealistic that crossing a single threshold would trigger both manual and saccadic execution. Similarly, recovered manual non-decision times are substantially longer (though not more variable) than expected motor execution durations for button presses. These patterns suggest that parts of what the model treats as non-decision time are likely decisional in nature, although perhaps related to "action decision" rather than the "value-based decision" of interest to the authors. To what extent these two processes neatly follow each other or overlap could be usefully considered.

      We have added a paragraph to the Discussion explaining how our model’s estimates of sensory and motor latencies relate to corresponding values inferred from physiology or behavioral manipulations (e.g., Bompas et al., 2024). Specifically, we write:

      “The key assumption of the PDG model is that there is a delay between the moment a choice is internally committed and the moment it is externally reported with a key press. Because eye movements are typically faster than manual responses (𝜏<sub>e</sub> < 𝜏<sub>m</sub> in our simulations), this delay creates a window during which gaze can already be directed toward the covertly chosen item before the response is formally registered. We do not interpret these non-decision latencies as irreducible physiological minima for moving the eyes or pressing a button (Bompas et al., 2025). Rather, they are inferred indirectly by fitting an additive non-decision-time parameter to the behavioral data, which we decompose into a sensory delay (𝜏<sub>s</sub>) and a manual execution delay (𝜏<sub>m</sub>). Values of 𝜏<sub>e</sub> are then chosen so that the model reproduces the observed magnitude of the behavioral effects. This estimation procedure has important limitations. Some participants show relatively “flat” chronometric functions: response times vary little with value despite otherwise normal psychometric performance. Such patterns likely reflect processes not explicitly represented in the model, including procrastination, reduced motivation, task-unrelated thought, or noise in item ratings. Within a drift-diffusion framework, however, these cases are accommodated by assigning a long non-decision time together with a short evidence-accumulation period (Table S1). Consequently, some estimated non-decision times are substantially longer than would be expected if they represented only sensory and motor delays. A further limitation is conceptual. We model non-decision time as occurring either before or after evidence accumulation, whereas in reality decisional and non-decisional components are likely temporally interleaved (Graziano et al., 2011). This simplification may also inflate the recovered latency estimates. With these caveats in mind, sensory and oculomotor delays on the order of 300 ms remain broadly plausible, although they likely lie near the upper end of a realistic range. The estimated eye-movement latency is especially long. For instance, in monkeys trained to report simple perceptual decisions with a saccade, roughly 100 ms elapses between the threshold-crossing signal in parietal cortex (or the superior colliculus) and the executed eye movement (Roitman and Shadlen, 2002; Stine et al., 2023). Crucially, however, varying the assumed non-decision latencies across a reasonable range does not alter the qualitative predictions of the model (Fig. 8).”

      Further, we have added a parameter sensitivity analysis. Importantly, although the magnitude of the predicted effects depend on the non-decision latencies, the qualitative aspect of these predictions do not (new Figure 8). Specifically, (i) the increasing tendency to look at the ultimately chosen item as time elapses (new Fig. 8A), (ii) the lack of an interaction between the last-fixation bias and overall value (Fig. 8B), and (iii) the absence of an effect of choice consistency on Δdwell (Fig. 8C) are all findings that are independent of 𝜏<sub>e</sub>.

      Reviewer #2 (Public review):

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Following this suggestion (and the reviewer’s elaboration in the private comments to the authors), we have substantially restructured the manuscript. Both aDDM predictions are now presented together (new Fig. 2), and Figs. 3–4 test these predictions across multiple food-choice datasets. In doing so, we no longer treat the data from Krajbich et al. (2010) separately, and we extend the analysis of the last-fixation–choice association (MELFB) to additional datasets. We note that the same datasets could not be used in both Figs. 3 and 4, as some lack information on the final fixation required for the MELFB analysis. Nevertheless, results are highly consistent across datasets and align with findings from a recent study by Ting & Gluth (2025), which independently identified and examined one of our key predictions; this work is now cited in the revised manuscript. Finally, to reduce redundancy, we have consolidated all aDDM variants and optimal models into a single figure (new Fig. 10).

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

      The key predictions of the model depend on the difference between the manual (𝜏<sub>m</sub>) and eye-movement-related (𝜏<sub>e</sub>) latencies. We have now added a parameter-sensitivity analysis to show how the model predictions depend on this difference. The new analysis shows that while the quantitative predictions do depend on the precise latency values, the results are qualitatively similar across values of 𝜏<sub>e</sub> (new Figure 8).

      Reviewer #3 (Public review):

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

      Thank you for this suggestion. We added a new paragraph to the discussion (paragraph #2), where we argue that it is sensible for a decision maker to direct the gaze to the chosen item once a covert choice commitment has been made, as the benefits of attending to a stimulus do not end with the decision itself. Specifically we now write:

      “Instead, these observations are better explained by a post-decision account of the gaze-choice association that is, one in which gaze shifts to the selected item after a covert commitment to a choice. We argue that directing gaze to the chosen item after a covert choice commitment is sensible, as the benefits of attending to a stimulus do not end with the decision itself. In naturalistic settings, for instance, selecting a food item is typically followed by the action of reaching toward it, where visual attention supports spatial localization and motor planning for the upcoming action. Although participants in our computerized task did not physically act on their choices, these sensorimotor processes are likely highly automatized and may still be engaged by default, even when not strictly required. Beyond motor preparation, post-decisional attention may also serve additional functions, such as facilitating sensory anticipation of the reward, supporting metacognitive evaluation of the decision, and contributing to value updating for future choices. From this perspective, a degree of attentional “stickiness” whereby the chosen item remains preferentially attended after commitment could emerge as an effectively optimal policy once these post-decisional processes are taken into account. Moreover, a specific feature of the task design may further reinforce this tendency: in the snacks paradigm, the unchosen item typically disappears from the screen immediately after a response is registered. It is therefore plausible that directing gaze to the chosen item after commitment partly reflects anticipation of the imminent disappearance of the unchosen option. To disentangle these mechanisms, it would be interesting for future work to test whether this attentional bias persists when the chosen item, rather than the unchosen one, is the stimulus that disappears upon response.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments:

      (1) Framing of the modelling approach

      The manuscript would benefit from acknowledging the known limitations of DDM-based frameworks, especially given that the entire study is conducted within these constraints. The introduction highlights successes of the DDM, but the manuscript does not mention any of its conceptual or empirical limitations.

      We are unsure about what specific limitations the reviewer has in mind, but we have added a paragraph to discussion mentioning some limitations, like the inflation of the non-decision times and the difficulty of interpreting the fit parameters (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (2) Dependence on non-decision time assumptions

      The alternative model's explanatory power appears to rely heavily on assumptions regarding the decomposition of non-decision time: fixed visual encoding (𝜏<sub>s</sub>= 0.3 s), manual non-decision time (𝜏<sub>m</sub>; two free parameters), and saccadic execution (𝜏<sub>e</sub>; fixed parameters μ<sub>e</sub> = 0.35, σ<sub>e</sub> = 0.11).

      - 𝜏<sub>e</sub> is substantially longer and more variable than typical saccadic execution times, suggesting it likely incorporates decisional components.

      - Estimated 𝜏<sub>m</sub> values are approximately twice as long as known manual execution durations.

      - σnd is more plausible, implying that variability is captured correctly but mean durations are not.

      Together, these points raise the possibility that portions of what the model treats as non-decision time are in fact part of a (action) decision process. Only then does it make sense to assume that Tm is usually larger than Te. If Tm and Te were truly execution delays, then Tm would always be larger than Te.

      You may find it helpful to consider the framework in Bompas et al. Psych Review (2024), which discusses in detail what non-decision time is likely to comprise across effectors.

      Thank you we have added (i) a sensitivity analysis showing that our results are robust to changes in the specific value used for the eye movement related latencies (new Fig. 8), and (ii) a new paragraph in Discussion addressing the issue of the mismatch between our parameter estimates and the manual and saccadic execution times (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (3) Code availability.

      The authors should consider sharing all relevant code and data publicly.

      We agree, we now share the code and data on GitHub and indicate so in the revised manuscript.

      Minor Comments:

      (1) Lines 74-77. These are not worded as predictions but as questions; one tests predictions, but answers questions. I feel it would be clearer to stick to predictions (like in the abstract), and the introduction could benefit from explaining these predictions in a bit more detail (I found it difficult to get my head around these predictions from the intro text only).

      We rewrote the section in the introduction where we provide a gist of the model predictions (last paragraph of Introduction). We agree with the reviewer that the previous explanation was not clear.

      (2) It is confusing that panel B appears to the left of panel A in Figure 2.

      We agree. We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed and the panels follow a more logical order.

      (3) Figure 3C - remove MATLAB toggles.

      Yes, thanks.

      (4) Figure 5A shows the proportion of left choices, but the text and legend refer to right choices.

      Good catch, thank you.

      Reviewer #2 (Recommendations for the authors):

      This may appear self-serving, but the authors seem to be unaware of some highly relevant work from our group. Most importantly, in a recent publication (Ting & Gluth, 2024, JEP General), we have already looked at the dependency of the last- (or final-) fixation bias on overall value in value-based (VB) and perceptual (P) decisions. In VB, we found a negative effect; in P we did not find a significant effect. This is largely consistent with the current results, showing a negative but not significant trend. Another relevant work is Gluth et al. (2020, Nat Hum Behav), where we extended the aDDM by assuming that the probability to fixate on an option is a function of the accumulated evidence for that option. It would be interesting to know whether this assumption changes the predictions of the aDDM. Finally, we just published a new theory on how people search for information to make efficient value-based decisions (Gluth et al., in press, Psychol Rev; https://osf.io/preprints/psyarxiv/3qzak_v2). Although this theory focuses on multi-attribute choices, it can be applied to "simple" choices, too (by assuming that there is only one attribute = value). Interestingly, while the model also mispredicts a (slight) increase of the last-fixation bias with overall value, it correctly predicts the independency of the dwell-time advantage effect on choice consistency as well as the small increase of the effect with RT (attached here is a figure to show this: [https://elife-rp.msubmit.net/elife-rp_files/2026/01/22/00149589/00/149589_0_attach_9_477122. pdf], and the match with the empirical data shown in Figure 3B and 12 is striking). In general, the model shares many features of the Callaway and Jang models, but does not need to assume a biased value prior, which the authors suggest is responsible for the misprediction of the second effect. I leave it up to the authors to discuss this new theory, but I wanted to point this out.

      Thank you for pointing this out; these are all relevant points and studies.

      We now note that the first of our predictions has recently been identified and tested by Ting and Gluth (2025).

      We also considered extending the manuscript with a variant of the model proposed by Gluth et al. (Psychological Review, 2026). In fact, we attempted to fit this model to the Krajbich et al. (2010) dataset under the assumption that the duration of each sampling epoch is a free parameter. We find this model very interesting. However, in our current implementation it appears to make the same qualitative prediction as the aDDM, namely that ΔDwell depends on choice consistency (see Author response image 1).

      Given this, we have decided not to include these results in the manuscript. It remains possible that with further development particularly with a more realistic specification of fixation durations (e.g., allowing them to depend on value) the model could account for the full set of observed effects. We think this would be best addressed in a separate study.

      That said, we do find the model promising, as it provides a better account than most of the alternative models we explored for the patterns shown in panels D, H, and I.

      Author response image 1.

      Fits of a variant of the MACS model (Gluth et al. 2026) to the data of Krajbich et al. (2010).

      The paper would benefit substantially from restructuring. The aDDM's predictions are provided first, together with the empirical data, and then the optimal models are discussed. But Figure 2 shows all of this together. Later, the new (PDG) model is elaborated, and its predictions are shown. Towards the end of the results, variations of the aDDM and combinations of aDDM and PDG are shown in a series of figures (8-11), followed by a last figure showing one of the tested effects in other datasets. All of this feels pretty much thrown together without a clear structure. For instance, the aDDM and the optimal models could be described together (or the optimal models get a separate figure). The additive variants could be described earlier. And some figures could be put into the supplement. And the empirical results of the different studies could be shown together.

      We fully agree with this suggestion. We have now restructured the manuscript along the lines proposed by the reviewer (see the more detailed explanation of the restructuring in our response to the public comments).

      I strongly suggest avoiding the term "influence" in the y-axis of Figure 2, upper row, as it implies causality. Similarly, in line 182, the term "causal influence" is used in the context of the Callaway model, but as far as I know, this is not what the model assumes.

      We replaced the y-axis label with “Association of last dwell with choice (β)”

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 2 - Panel labels for A and B are reversed?

      We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed.

      (2) Does 3C include a .pdf screenshot?

      Thank you, it’s a Matlab bug on Mac. I guess they want us to switch to Python -:)

      (3) Figure 4 - It would be helpful if the green line were defined in the figure legend.

      Added

      (4) The effect size in 5B looks much more dramatic than in 2B(A?) - Is this for one example subject as opposed to all subjects? Please clarify what is different about the data.

      We are no longer showing the psychometric functions in Figure 2.

      (5) Line 252 - they say they compared the probability of choosing the right item (Fig. 5B) by the y-labels of that figure, which are all p(choose left).

      Yes, corrected now.

      (6) In general, they reference the subpanels of Figure 5 out of order, which causes the reader to jump around. They might consider reordering the panels of the figure so they follow the ordering of descriptions in the text.

      We agree, we have rearranged the figure panels to follow the ordering of the descriptions in the text.

    1. eLife Assessment

      This important study measures single-unit activity in area MT of awake-behaving monkeys to test the idea that sensory adaptation contributes to flexible evidence accumulation during decision making. The authors provide compelling evidence that adaptation to different temporal contexts shapes both perceptual judgements and neural responses. Although the precise computational mechanisms underlying these effects remain uncertain, the results support the conclusion that recent sensory history influences the temporal dynamics of decision formation. This work will be of interest to researchers studying visual perception, sensory adaptation, and decision making.

    2. Reviewer #1 (Public review):

      McGaughey and Gold ask where in the decision process the flexibility of evidence accumulation arises, proposing that it is not solely a property of downstream integrators but is also supported by stimulus-specific sensory adaptation in the middle temporal area (MT). Recording single-unit activity in rhesus macaques during a motion direction-discrimination task in which an adapting stimulus of varying temporal stability precedes an identical test stimulus, they find that more rapidly changing contexts produce weaker and less discriminable MT responses to the test stimulus, which they argue accounts in part for context-dependent changes in decision-making behavior. Through session-level correlations they further identify pupil-linked arousal as a parallel, apparently separable contributor.

      The main strength is the shift of perspective toward the encoding stage: rather than treating MT as a static input to flexible downstream integrators, the authors show that early sensory cortex can itself contribute adaptive, context-dependent signals that shape behavior. The conceptual advance is supported by a well-designed paradigm-total exposure to each motion direction is matched across conditions and the test stimulus is held identical-together with single-unit recordings and simultaneous pupillometry. The behavioral effect is consistent across three animals, and the fact that context-dependent differences emerge over repeated stimulus presentations within a trial, rather than as a sustained baseline offset across blocks, ties the effect convincingly to stimulus-specific adaptation.

      The behavioral effect constrains the temporal dynamics of decision formation but does not uniquely identify its algorithmic basis: a leak, a saturating non-linearity, or a reduction in the gain of integration are all compatible with a shallower rise of accuracy with viewing time, and the reduced MT discriminability is itself an encoding-stage efficiency effect of this kind. The manuscript appropriately treats the algorithmic basis as unresolved, noting that distinguishing these accounts would require analyses not available here, such as reverse-correlation or motion-energy kernels with lower-coherence test stimuli.

      The inference that the adaptation- and arousal-related signals operate independently rests on the absence of session-wise correlations between the neural and pupil measures and their behavioral contributions. Given the noise in the trial-wise estimates, this is best read as consistent with, rather than demonstrating, true independence, as the authors note.

      Overall, the authors largely achieve their aim of showing that sensory adaptation in MT shapes the evidence available for time-dependent perceptual decisions. The evidence for a sensory-encoding contribution is convincing, while the claim of independence between adaptation and arousal is more tentative and is framed as such.

    3. Reviewer #2 (Public review):

      McGaughey and Gold trained rhesus macaque monkeys to perform a motion-direction discrimination task in which a behaviorally irrelevant adapting stimulus with either fast or slow direction alternations preceded a variable-duration test stimulus, while simultaneously recording single-unit activity in area MT and pupil diameter. They report that adaptation to the more rapidly changing stimulus was associated with reduced behavioral sensitivity, attenuated test-evoked MT responses, and larger pupil-linked arousal signals. The authors interpret these behavioral changes as evidence for context-dependent adjustments to the temporal dynamics of decision formation and argue that these adjustments are supported by both sensory adaptation in MT and arousal-related mechanisms. More broadly, they conclude that flexible evidence accumulation in dynamic environments arises from distributed adjustments across sensory encoding and neuromodulatory systems rather than solely from changes within a downstream accumulator. If correct, this interpretation has important implications not only for our understanding of perceptual decision making, but also for broader theories concerning the functional role of sensory adaptation.

      The conclusions of the paper are generally supported by the data. Evidence for adaptation-induced changes in sensory encoding, behavior, and pupil dynamics is convincing, and the revised manuscript substantially strengthens the connection between the behavioral findings and the proposed decision-making framework.

      Comments on revised version.

      The revised manuscript provides a clearer account of how recent stimulus history influences behavioral performance. In the original version, aspects of the psychometric functions were interpreted as evidence for a more leaky evidence-accumulation process, although some of these effects could potentially have reflected alternative mechanisms, including influences of the adapting stimulus on short-duration trials. The additional analyses and discussion included in the revision clarify that information from the adapting stimulus contributes to behavior at short viewing durations and appropriately temper claims regarding the specific computational mechanism underlying the observed behavioral effects. While the data do not uniquely identify whether these effects arise from changes in leak, other nonlinearities, or related decision processes, they provide convincing evidence that recent temporal context influences the temporal dynamics of decision formation.

      My original review also noted that different sections of the manuscript relied on different behavioral metrics and analytical approaches when relating behavioral changes to neural and pupil-linked measures. The revised manuscript now provides a clearer rationale for these choices, including distinctions arising from the different trial types and time windows used in the neural and pupil analyses.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) Alternative mechanisms for performance differences.

      The authors assume that the difference in performance between the low-switch (LS) and high-switch (HS) frequency conditions is explained by a change in the "leakiness" of integration. However, several other mechanisms could potentially explain this effect:

      (1) Temporal Uncertainty: Integration might start later in the HS condition, leading to lower performance.

      (2) Reduced Efficiency: Integration could be less efficient in the HS condition (i.e., lower signal-to-noise ratio) without a change in the leak parameter itself.

      (3) Evidence Contamination: Motion information from the adapting stimulus in the HS condition may be integrated rather than ignored, which might be the case since the transition from the adapting to the test stimulus is not externally cued.

      To distinguish between these alternatives, I suggest two possible analyses. First, a formal model comparison could be performed, though I acknowledge this may be inconclusive in the absence of response-time data. Second, an analysis of motion energy kernels could be revealing; the leak hypothesis makes the specific prediction that for long test stimuli, early samples should contribute more to the choice in the LS condition than in the HS condition, relative to late samples.

      We thank the reviewer for raising these important points. We agree that we cannot definitively identify the algorithmic underpinnings of the behavioral effects we report and have made substantial revisions to the manuscript to be clearer about what is supported and what is speculative in our claims. Most importantly, we agree that we do not know if the context-dependent differences in how accuracy depends on viewing time are based on adjustments to a leak or to something else (e.g., a saturating non-linearity, as we identified in Glaze et al, 2015, that is separate from the leak itself), which we cannot resolve with this dataset, even with more formal model comparisons. We therefore:

      Changed the wording throughout the manuscript to refer to changes in leakiness as just one of several possible sources of the behavioral differences. We also added this point to the list of “limitations” (and possible future directions, including using motion-energy kernels, which would require us to use lower-coherence test stimuli) in the Discussion (L487-493).

      Added a new figure panel (Fig. 2D), a new Extended Data figure (Extended Data Fig. 3), and additional explanatory text (L168-175) that collectively describe the behavior in more detail, including quantifying a “crossover” dynamic similar to what we reported previously (Glaze et al, 2015).

      Added new explanations (L152-163) and analyses (Extended Data Fig. 9) indicating that the monkeys used some information from the end of the adapting stimulus to inform their decisions, which accounts for the patterns of choices at the shortest viewing durations.

      Indicate that the context-dependent differences in the slopes of the psychometric functions (and complementary analyses based on “raw” accuracy measures as a function of binned viewing duration) rule out the temporal uncertainty and evidence contamination explanations, but are consistent with effects on the temporal dynamics of the decision process (L175-179).

      (2) Independence of neural and pupil-linked signals.

      The authors take the lack of session-wise correlation between context-dependent contributions from neural and pupil terms as evidence that these two signals provide independent contributions to the behavioral effect. However, could this lack of correlation simply be a result of high variability or noise in these estimates? The data shown in Figure 7B suggests that measurements are very noisy, which might obscure a potential relationship.

      We agree that the lack of session-wise correlation between neural and pupil terms cannot be taken as definitive evidence of independence. We have both softened the language around the claim (L368) and added a sentence to the Discussion (L464-468) acknowledging that this lack of correlation may reflect underlying noise and/or variability rather than true independence of the underlying mechanisms.

      Reviewer #1 (Recommendations for the authors):

      (3) The neural data analyses rely fundamentally on "switch" trials (Figures 3-5). It might be informative to also examine "non-switch" trials to see if there are specific neural markers indicating the exact moment the motion stimulus becomes behaviorally relevant. Given that this may fall outside the primary focus of the paper, it is up to the authors whether to pursue this line of inquiry.

      We thank the reviewer for this suggestion. We agree and have added new analyses of data from non-switch trials (Extended Data Fig. 9), which show some effects of stimulus information from the adapting epoch on the monkeys’ choices, as we detail below in response to related comments from the other reviewers.

      Reviewer #2 (Public review):

      Aspects of the behavioral analysis would benefit from a tighter connection between theoretical claims about evidence accumulation and the empirical features of the psychometric functions. For example, the rightward shifts observed across adapting conditions are interpreted as consistent with a reset of accumulation on switch trials, but similar patterns could also arise from failures to detect the test stimulus on a subset of trials, leading responses to default to the final adaptor direction. Likewise, changes in psychometric slope and asymptote are attributed to differences in evidence accumulation without explicit modelling or consideration of alternative explanations.

      Clarifying how specific features of the psychometric functions map onto distinct components of the decision process will strengthen the link between the theoretical framework and the behavioral data.

      We agree and have made substantial revisions to address these important points. Specifically, we added a new figure panel (Fig. 2D), new Extended Data Figures (3 and 9), and several lines of explanatory text (L152-179) that collectively describe the behavior in more detail, including clarifying that: 1) for the shortest viewing durations, the monkeys’ decisions were informed by information from the adapting stimulus, which accounts for generally lower accuracy on LSF (longer exposure to the final adapting direction, thus more accumulated evidence for that direction before processing the switch) vs. HSF (shorter exposure to the final adapting direction, thus less accumulated evidence for that direction before processing the switch) switch trials; and 2) as viewing duration increased, the rate of rise of accuracy versus viewing duration was higher for LSF vs. HSF trials, implying differences in the process of evidence accumulation. As detailed in our response to a similar comment from Reviewer 1, above, we are now careful to temper our claims about the specific computational basis (e.g., a leak or other form of nonlinearity) for these differences.

      We also de-emphasized our treatment of the asymptotes of the psychometric functions. In principle, these regimes could give insights into leakiness (which can limit the total amount of information that can be accumulated) and lapses (which are measured at the asymptotes). In practice, however, the long-duration trials that constitute the asymptotes were relatively under sampled (to promote the unpredictability of the offset of the stimulus, which we believed was the more important consideration when designing the experiment), yielding unreliable estimates.

      A slight concern is the lack of a consistent analytical approach for relating behavioral changes to neural and pupil-linked measures. Different sections of the manuscript rely on different behavioral metrics-such as differences in accuracy within a selected stimulus-duration range (e.g., Figure 5C) or psychometric slope differences (Figure 6C) without clear justification for these choices. The analytical approach likewise varies between simple correlational analyses (Figure 5C, Figure 6C), pseudo-experimental group comparisons (Figures 5D, E), and the inclusion of neural or pupil terms in the behavioral psychometric regression model (Figure 7B). While each metric and approach may be defensible in isolation, adopting a more consistent framework will help convince readers that the reported effects are robust and not contingent on the selective choice of metric or analysis.

      We thank the reviewer for this thoughtful critique and agree that the rationale for our choice of behavioral metrics and analytical approaches could be stated more clearly. We have added text to the relevant sections of the Results (L247-251) clarifying these choices. In particular:

      The neural analyses (Figures 3D-E, Figure 4, Figure 5D-E) focused on preferred-motion switch trials, because: 1) low switch-frequency non-switch trials provide an additional 800 ms of exposure to the final adapting-stimulus motion direction relative to high switch-frequency non-switch trials, which confounds comparisons of context-dependent evidence encoding between conditions, and 2) MT neurons exhibit minimal responses to null motion (although note that we also included analyses based on ROC area, which is computed from both preferred- and null-motion switch trials, to account for possible contributions of null-motion responses; Figure 5A-C). Thus, to ensure a meaningful comparison between neural and behavioral measures, we used behavioral accuracy on switch trials as the relevant metric in Figure 5C-E, rather than psychometric slope, which is estimated across both switch and non-switch trials.

      The pupil analyses (Figure 6) focused on a time window preceding test-stimulus onset, representing the arousal state around when the decision process started, and included both switch and non-switch trials. Thus, for these analyses we used psychometric slope, which is estimated across both switch and non-switch trials.

      We used several different analyses to compare and contrast the neural-behavioral and pupil-behavioral relationships because they provide complementary and useful insights. The correlational analyses in Figures 5C and 6C characterize session-level relationships between neural/pupil signals and behavior. The group comparisons in Figures 5D–E provide a complementary visualization of the same relationship. The model-based approach in Figure 7 then allows direct quantification of the trial-wise contributions of each signal to behavior within a common framework. Importantly, the conclusions drawn from each approach converge on the same interpretation, which we believe speaks to the robustness of the reported effects.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 2 legend. Description of 'running average (5-trial window)' is unclear - presumably this is a running average in stimulus space rather than across trials.

      We thank the reviewer for flagging this ambiguity. We have updated the legend (L136-137) to clarify that the running average is computed across trials sorted by test-stimulus duration.

      (2) L158. Difficult to establish an asymptotic performance level for HSF conditions within the stimulus duration range tested.

      We have removed the reference to asymptotic performance and replaced it with a discussion of performance on longer-duration switch trials in the context of the newly added Figure 2D.

      (3) L515 Equation 1. While this is a standard formulation of lapse rate in psychometric functions, the construction here in terms of switch probability is not standard. Given the task and training, it seems more likely that on lapse trials, the animal will respond according to the last adapted direction (rather than randomly switch/stay with equal probability).

      We thank the reviewer for this point. We agree that it is possible that on at least some of the “lapse” trials the monkeys may respond according to the final adapting-stimulus direction rather than choosing randomly. However, we cannot distinguish those alternatives using this task design. We include a statement to this effect in Methods (L569-571).

      To explore the idea further, we refit the behavioral data using separate upper and lower asymptotes corresponding to lapse rates on switch and non-switch trials, respectively. Across monkeys, there were no significant differences between upper and lower lapse rates for either low (Wilcoxon signed-rank test for equal medians: p = 0.15, Cohen's d = -0.13) or high switchfrequency (p = 0.07, Cohen's d = -0.16) conditions. So, at the very least, there was no evidence for lapse-like errors driven by switch- (or non-switch-) specific defaults to the final adapting direction.

      (4) L256. Statistical significance of attenuation is not directly tested here.

      We have replaced "were attenuated" with "we did not identify any reliable context-stability differences" (L297) to accurately reflect what was directly tested without implying a statistical comparison between groups of sessions that was not performed.

      (5) L429. Does the increase in explanatory power warrant the increased complexity of the model here?

      We thank the reviewer for raising this important point. We used Tjur's pseudo-R<sup>2</sup> because it does not increase by default with added model complexity, making it more conservative than other R<sup>2</sup> measures in this respect. Tjur's pseudo-R<sup>2</sup> is a coefficient of discrimination, and as such its value increases only when additional terms improve the model's ability to separate predicted probabilities across response outcomes. Thus, the observed increases in explanatory power when adding neural or pupil terms reflect real improvements in discriminability rather than an artifact of model complexity. We have added a brief clarification of this point to the Methods (L662-664).

      Reviewer #3 (Public review):

      The task design may not be optimal. While the amount of time the monkey is exposed to each motion direction during the adapting stimulus is matched, it's hard to know if the reduced MT responses to the test stimulus are truly due to the greater frequency of switches during the HSF adapting stimulus or because the monkeys have been exposed to more repetitions of the stimulus. It's increased sensory adaptation in either case, but it makes it problematic to interpret this as temporal context-dependent adaptation specifically. I think this could potentially be partially addressed by an analysis that is in the paper, but could potentially be emphasized/fleshed out more, specifically the results shown in Figure 4D that seem to show that most of the reduction in neural response for adapting units occurs between the first and second stimuli.

      The reviewer raises an important point. The number of stimulus repetitions and switch frequency are confounded in the experimental design, making it difficult to attribute context-dependent differences in MT responses to the temporal pattern of switches rather than to accumulated repetitions. We also note, as the reviewer acknowledges, the observed differences reflect sensory adaptation either way. Figure 4D does offer relevant evidence, suggesting that a majority of the change in neural response occurred with just one stimulus repetition. This finding complicates an interpretation where adaptation scales with the number of stimulus repetitions. We have added several lines to the Results about these points (L231-233).

      The pupillometric analysis seems to be an indirect way of assessing whether the accumulator itself might be modulated by temporal context, but the link could be made clearer. The authors show that context-dependent behavior is related to pupil size, which is related to arousal/neuromodulation, but it would be helpful to have some idea of what neural mechanisms underlying adaptive decision-making are actually impacted by this neuromodulation. Lacking neural data to address this question (e.g., from a brain region proposed to be involved in the accumulation process), at least more discussion of this would be helpful. Essentially, I'm unsure of how to interpret the pupil results: the argument that temporal context affects instantaneous evidence encoding in MT that then drives the accumulator is very clear, but I am a bit confused about what, mechanistically, I should think about the effect of neuromodulation doing.

      We thank the reviewer for this thoughtful comment and agree that the mechanistic interpretation of the pupil results could be made clearer. We acknowledge that we cannot directly identify the neural mechanisms underlying the arousal-related contributions to adaptive evidence accumulation from pupil data alone, given that pupil size is an indirect and imperfect proxy for neural (e.g., LC-NE system) activity. However, we can offer some informed conjecture and have added to the Discussion (L469-482) in an effort to elaborate on possible mechanisms.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract could be retooled - does not emphasize the pupillometry/arousal results very much, and they are presented more as a control than an independent result.

      We agree and have revised the Abstract accordingly.

      (2) Do all neural/pupil analyses use only switch trials? Sometimes the figure captions do specify only switch trials, but not everywhere. It would be helpful to specify either in the Methods or at the beginning of each figure caption that all subplots show switch trial results. Also, if you do always use switch trials, it would be useful to see in the Supplement how the non-switch trial results differ from switch trials. It seems like they may in interesting ways based on the behavioral results (supporting a reset of evidence accumulation on switch but not non-switch trials).

      We thank the reviewer for flagging these important points. We have added a justification for switch trials (L186-190) as well as clarification about which trial types were used for which analyses (L246-249) and information about trial types to relevant figure captions. We have also added a new Extended Data figure (Extended Data Fig. 9) examining relationships between neural activity and behavior on non-switch trials. As inferred by the reviewer, behavior on non-switch trials is consistent with the use of information from the adapting stimulus.

      (3) In Figure 3C, 5B, etc, when computing firing rate for the test stimulus (50-500 ms), are differently sized windows used to compute the rate for different test stimulus durations (since some will be <500 ms)? Or are only trials where the test stimulus duration is > 500 ms used for this analysis?

      We thank the reviewer for raising this point. To clarify, the 50–500 ms window does not reflect a fixed window applicable for all trials. Rather, neural activity from 50 ms after test-stimulus onset through test-stimulus offset was included for each trial, with 500 ms serving as the upper bound for trials with longer durations (> 500 ms). We have clarified this in the Methods (L607-610) to avoid ambiguity.

      (4) I think it might be better to be consistent with the time windows used for analysis; specifically, to choose either the 50-500 ms window used in Figures 3, 4, and 5B, or the 200- 400 ms window used for the remaining analyses in Figure 5.

      We agree that using the same window for all of the analyses would improve consistency, but not doing so provides advantages that we believe take precedent and now describe in more detail. The broader 50–500 ms window used for Figures 3, 4, and 5B was chosen to characterize MT neural activity over a relatively large a time window, ensuring that every trial contributes to each estimate. Because test-stimulus durations were drawn from a truncated exponential distribution (100–1200 ms), restricting these analyses to the 200–400 ms window would have excluded the substantial proportion of trials with durations <200 ms (but would yield similar figures and conclusions). The narrower window used in subsequent analyses allows us to focus on the conditions that exhibited the biggest modulations of neural activity when comparing them to behavior.

      (5) Similarly, provide justification for using only trials ending 375-600 ms after test stimulus onset for the behavioral correlations. It seems reasonable to choose a subset of test stimulus durations where the monkeys' behavior is greater than chance but less than ceiling, but it would be good to specify this so that it doesn't seem arbitrary.

      We agree and have added text to make this important point (L249-251).

    1. eLife Assessment

      By investigating spine nanostructure and dynamics across multiple genetic mouse models for neurodevelopmental disorders, this important study has the potential to uncover convergent or divergent synaptic phenotypes that may be specifically associated with autism versus schizophrenia risk. The imaging and overall breadth of the methods are convincing. The purely in vitro nature of the study slightly limits the generalisability of the findings, though these limitations are acknowledged and discussed in the manuscript.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Kashiwagi et al. undertook a population analysis of dendritic spine nanostructure applied to the objective grouping of 8 mouse models of neuropsychiatric disorders. They report that spine morphology in cultured hippocampal neurons shows a higher similarity among schizophrenia mouse models (compared with autism spectrum disorder (ASD) mouse models) and identify an effect of Ecrg4 (encoding small secretory peptides) on spine dynamics and shape in these models.

      Strengths:

      The study developed a method for objectively comparing spine properties in primary hippocampal neuron cultures from 8 mouse models of psychiatric disorders at the population level using high-resolution structured illumination microscopy (SIM) imaging. This novel technique identified two distinct groups of mouse models according to the population-level spine properties: those with ASD-related gene mutations and those with schizophrenia-related gene mutations. Functional studies, including gene knockdown and overexpression experiments, identified an effect of Ecrg4 on the spine phenotype of the schizophrenia model mice.

      Weaknesses:

      The main weakness is that the study is wholly in vitro, using cultured hippocampal neurons. The authors present this as an advantage, however, arguing that spine morphology as measured in a reduced culture system can demonstrate direct effects of gene mutations on neuronal phenotypes in the absence of indirect influences from nonneuronal cells or specific environments.

    3. Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for timelapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. My major concerns from the initial manuscript, especially regarding cherry picking and circularity have been addressed with revised analytical approaches. I have some remaining minor comments.

      (1) The comparison between two wild-type samples versus wild-type-mutant samples is helpful - I think this could be added to the manuscript.

      As suggested, we added the figure comparing two wild-type samples against wild-type mutant samples as Supplementary Figure 2. 

      (2) For results of time-lapse imaging - please spell out in the results section the direction of change (lines 270 - 277).

      As suggested, we added the direction of change (an increase in the turnover rate) to the text (page 12, lines 270-271).

      (3) Using linear mixed effect models for statistical analysis is a significant improvement. While a sample size (n) of mice = 3 is not ideal, I think given the multiple different mouse lines used and intensity of analysis, this is probably the best that can be done, although further validation in larger samples eventually is to be hoped for.

      We appreciate the reviewer for recognizing the effort required to collect data across multiple mouse lines.

      (4) The revised text is much improved, but I still think the authors should be upfront somewhere in the text that the schizophrenia-associated genes can only confer biased risk for schizophrenia (and that the clinical phenotype can also include autism). As I said before, I think this is the best we can do and I agree with their choices, but it is important not to overstate the link. The differences they see make it clear that these are still relevant distinctions.

      As suggested by the reviewer, we further modified the discussion related to the comparison between ASD- and schizophrenia-associated mouse models (pages 23-24, lines 508-522).

      “The nanoscale features of dendritic spines in mouse models of Nlgn3<sup>R451C/(y or R451C)</sup>, Syngap1<sup>+/−</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, which we classified as being related to ASD, are highly heterogeneous. This heterogeneity may reflect the broad clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. Accordingly, these four mouse models may represent distinct subgroups characterized by different degrees or forms of hippocampal dysfunction. Notably, among the ASD-related models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those found in the 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> mouse models. Although we classified 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> as schizophrenia-related models, both 22q11.2 deletion syndrome and Setd1a haploinsufficiency in humans are also associated with ASD, suggesting substantial overlap in the genetic risk factors underlying ASD and schizophrenia. Further systematic analyses linking rare genetic variants to synaptic phenotypes in mouse models may provide important insights into the mechanisms underlying both shared and disorder-specific synaptic alterations in neurodevelopmental and psychiatric disorders.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I would suggest that it might be preferable to use the word 'neuropsychiatric' rather than 'mental' in the title.

      As suggested, we modified the manuscript title.

      (2) I think it would be clearer to say that DEGs are listed if present 'in three or more models' rather than >2 (I appreciate the latter is mathematically clear, but can easily be read as 2 or more if reading fast). This is changed in the figure legend, but I suggest it is also changed in the main text (line 352-3)

      As suggested, we changed the main text to incorporate "in three or more models" (page 16, line 352).

      (3) Please add to Methods (line 557) that 'control cultures were prepared from littermate embryos....'

      As suggested, we added the phrase "control cultures were prepared from littermate embryos" (page 26, line 559).

      (4) Sorry to add something, but please could the authors add a definition of how they calculate spine turnover (and add units to the y axis of Figure 5A-C)?

      As suggested, we modified the y-axis of Figure 5A-C (% as unit) and added the method of calculating spine turnover rate in the text (page 36, lines 808-811).

    1. eLife Assessment

      This useful study combines experiments and mathematical modeling to show that antibiotic protection provided by resistant cells can extend across both surface-associated and freely growing bacterial populations. Notably, they show that treatment efficacy depends on population composition and density. The evidence supporting the main conclusions is incomplete, primarily because the biofilm context is not adequately characterized and demonstrated, raising the concern that it might represent only an aggregate of cells on the surface (rather than a biofilm) under the studied experimental conditions.

    2. Reviewer #1 (Public review):

      Summary:

      This important study examines how antibiotic-resistant bacterial cells can protect neighboring sensitive cells in mixed populations that occupy both surface-associated and freely growing states. Using experiments in Enterococcus faecalis together with a mathematical model, the authors test the hypothesis that protection would be stronger in biofilm-associated populations, but instead find that resistance-mediated protection extends broadly across both population types. The work provides evidence that antibiotic efficacy depends strongly on community composition, population density, and density-dependent detoxification dynamics.

      Strengths:

      A major strength of the study is the close integration of experimental measurements with a relatively simple quantitative model that captures many of the observed population dynamics. In particular, the work highlights how interactions between antibiotic detoxification, cellular growth, and saturation at carrying capacity can generate nonintuitive behavior, including the reported population inversion effect. The agreement between the well-mixed model and the experimental observations is convincing, and the spatial analyses suggest that cells within the biofilm are sufficiently intermixed that large-scale spatial segregation is unlikely to dominate the observed behavior.

      Weaknesses:

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The study also raises interesting questions regarding the role of spatial structure and exchange between planktonic and biofilm-associated populations. It would be informative to explore whether biofilm-specific protection becomes more pronounced at lower antibiotic concentrations, where local detoxification may compete more directly with antibiotic penetration into the biofilm, and in this context, the dynamics of exchange between biofilm and planktonic populations would be interesting to understand. Overall, the evidence supporting the central conclusions is convincing, and the study will likely be of broad interest to researchers studying microbial communities, antibiotic resistance, and collective population dynamics.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Martins et al. examined the cooperative response of E. faecalis cells to beta-lactams, in both planktonic culture and in biofilm. They found that the competition outcome between the susceptible and resistant strains is frequency dependent; they have also quantified how the competition curves change with inoculation OD and antibiotic concentration. To the authors' surprise, the competition dynamics are not that different in biofilm and in planktonic culture, which the author attributed to the unstructured nature of the thus-grown E. faecalis biofilms, quantified through correlation analysis. Using a well-mixed model capturing growth, death, and drug degradation by the resistant cells, the authors were able to quantitatively capture the experimental observation.

      Strengths:

      Overall, the data presented are solid. Although there is not much surprise after the understanding that the E. faecalis biofilm is unstructured, the manuscript still provides a useful "null case", so to speak, for researchers in the field when considering antibiotics in the context of biofilm. The theoretical model presented and the procedure of fitting the experimental data are useful to the research community.

      Weaknesses:

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

    4. Reviewer #3 (Public review):

      Summary:

      The authors studied social aspects of antibiotic resistance by co-cultivating antibiotic-resistant and sensitive Enterococcus faecalis (an important pathogen) as biofilms to assess the extent to which sensitive cells can take advantage of the protection provided by resistant cells against both a beta-lactam antibiotic and in the presence of a B-lacatamase inhibitor. By quantifying the proportion of each cell type using fluorescence microscopy, they conclude that protection is provided equally in the biofilm and planktonically, and that the biofilm is completely unstructured with regard to the locations of the two cell types. A mathematical model is then used to show that no spatial information is needed to recapitulate the results and that the protective effect can be described completely by the growth rates of the two cell types and the affinity of the β-lactamase to the antibiotic and inhibitor. The strength of evidence is difficult to assess due to unclear descriptions of some methods, and the significance of the findings is limited by the experimental setup, where antibiotics were added very close to the time of inoculation.

      Strengths:

      The co-cultivation of antibiotic-resistant and sensitive bacteria allows for exploration of the social aspects of antibiotic resistance. Fluorescently-tagged strains allow for unambiguous tracking of the two cell types. The simultaneous analysis of biofilm and planktonic cells enables insight into whether these different growth modalities are influenced by social aspects of antibiotic resistance. In analyzing the structure of the biofilm, the use of a null model with randomized cell positions allows for an accurate determination of whether the observed data are due to some effect; however, as noted below, there is a caveat to this analysis. The broad observation that biofilm and planktonic populations are linked is generally supported by the data; however, this result is closely tied to the experimental setup used. The development of a mathematical model that can recapitulate results from a second set of data with values obtained from fitting a different set of data shows robustness of the model for using it to explain the results.

      Weaknesses:

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. The described 'population inversion' effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed. Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account. The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case. The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

    5. Author response:

      We would like to thank the editors for their interest in our work and the three referees for their time and careful reading of the manuscript. The reviewers have provided a series of helpful suggestions that we discuss in this provisional reply and will seek to address in the revised version of the manuscript.

      The main concern raised is that the bacterial community we refer to as a biofilm may instead correspond to a cell aggregate. Following the passing of Prof. Kevin Wood, in whose lab the experimental work was carried out, our ability to perform additional experiments is limited. Nevertheless, we plan to wash and fluorescently stain the extracellular matrix before imaging to measure the extent to which the observed bacterial community is an attached biofilm. In the meantime, we would like to highlight the work of Wen Yu et al. [1], in which E. faecalis biofilms were grown in 96-well plates under antibiotic stress. In particular, one of the strains of E. faecalis used in this article was OG1RF, the same strain used in our study. Crystal violet staining was used to quantify biofilm biomass, providing evidence for biofilm formation under those conditions. While we recognize that the experimental setup differs from ours and that the OG1RF sample used did not contain fluorescent and resistance plasmids, these results nevertheless support the expectation that OG1RF will readily form biofilms.

      Reviewer #1 (Public review):

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The timescale of drug degradation is an important system metric. We appreciate the referee’s suggestion to quantify antibiotic concentration over time. We plan to perform experiments in which samples are collected from the culture at fixed time intervals. After removing the bacteria from the samples via centrifugation, serial dilutions of the supernatant will then be spotted on a lawn of sensitive cells to measure the antibiotic efficacy at each time point.

      Reviewer #2 (Public review):

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

      We agree with the referee that additional evidence would strengthen our study. As noted above, we will perform additional experiments to demonstrate the presence of an attached biofilm.

      Reviewer #3 (Public review):

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. [...] The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

      The experiment was designed to address how coupled planktonic and biofilm populations develop in the presence of antibiotics, which we will more explicitly discuss in the revised manuscript. We do agree that investigating how mature biofilms and their planktonic populations respond to antibiotic stress is an exciting direction for future studies. However, we believe that is beyond the scope of the our study on the development of coupled populations. We will be sure to explicitly identify this limitation in our revisions.

      The described ‘population inversion’ effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed.

      The reviewer is correct that this ‘population inversion’ is a frequency-dependent (perhaps also density-dependent) effect, and we should have situated it within that broader ecological framework. We will use this terminology in our revisions. We do want to acknowledge that the late Dr. Kevin Wood was fond of this phrasing to describe the reversal of the dominant strain, which is not necessarily true for frequency-dependent effects. Although we do not know for certain, we suspect this was a play on the ‘population inversion’ term used in quantum physics, used to describe a system in which its excited state (high energy) population unexpectedly outnumbers its ground state (low energy) population.

      The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case.

      If the reviewer is indeed discussing Figures 1a and 1b, these are not comparable as the starting fractions differ. On the other hand, if the reviewer was talking about Figures 2a and 2b (which is more clearly discussed by looking at Figures 2c and 2f), we agree that it appears that planktonic communities tend to have a slightly greater frequency of sensitive cells than the biofilms. We will be sure to highlight this observation and possible explanations in our revisions. However, given the uncertainty in these observations, we do not believe the differences are sufficient to alter our overall conclusion that final resistant fractions in biofilm and planktonic populations are quantitatively similar. Furthermore, the no drug treatment shows the same trend, which suggests it’s an effect of different growth dynamics of these two strains at high density rather than driven by the protective effects of resistance cells.

      Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account.

      The entirety of the Z stack was used to measure the final resistant fraction in the biofilm. On the other hand, we used only the densest slice of the Z stack to calculate the correlations. The correlations follow the same trend when calculated over less dense slices, but as density decreases, noise increases, so such plots did not bring more clarity to our conclusions and were not included in the manuscript. Additionally, only horizontal correlations (over a slice) were calculated because consecutive Z-stack slices were imaged with a 2.5 µm spacing. Given that the average cell diameter is approximately 1 µm, calculating vertical correlations may miss neighboring cells located between imaged slices, making such measurements unreliable. We will clarify the points raised by the reviewer in the results section and add more detail to the imaging methods section in the revised manuscript.

      References

      (1) Wen Yu, Kelsey M. Hallinen, and Kevin B. Wood. “Interplay between Antibiotic Efficacy and Drug-Induced Lysis Underlies Enhanced Biofilm Formation at Subinhibitory Drug Concentrations”. In: Antimicrobial Agents and Chemotherapy 62.1 (Dec. 2017), 10.1128/aac.01603–17. doi: 10.1128/aac.01603-17. url: https://journals.asm.org/doi/10.1128/aac.0160317 (visited on 01/11/2026).

    1. eLife Assessment

      This important study provides the first in vivo evidence that nonsense-mediated mRNA decay (NMD) in mature astrocytes regulates astrocyte function, synaptic plasticity, and anxiety-related behavior. Using a broad range of approaches, the authors show that conditional deletion of Upf2 alters astrocyte morphology and calcium signaling while impairing synaptic transmission and plasticity, providing solid support for the central conclusion that astrocytic NMD influences neural circuit function. Some key mechanistic claims remain incompletely supported, including whether phenotypes reflect astrocyte remodeling versus loss, the interpretation of synaptic engulfment data, the link between NMD targets and calcium signaling, and the extent to which calcium dysregulation explains the observed synaptic and behavioral effects.

    2. Reviewer #1 (Public review):

      Summary:

      Lituma and colleagues investigate the role of NMD in astrocytes, an underexplored question given that prior work on NMD in the brain has focused exclusively on neurons. Using a tamoxifen-inducible, astrocyte-specific Upf2 conditional knockout (cKO) mouse, they report that loss of astrocytic NMD causes: (1) reductions in astrocyte cell volume and surface area across hippocampus, visual cortex, and prefrontal cortex; (2) decreased excitatory synapse density, reduced dendritic spine density, and impaired synaptic engulfment; (3) deficits in basal synaptic transmission and LTP, with selective impairment of mGluR-LTD; (4) elevated spontaneous calcium transients in astrocytes; and (5) anxiety-like behavior in the elevated plus maze (EPM) and contextual fear conditioning paradigms. Transcriptomic analysis of FACS-isolated astrocytes identifies 277 differentially expressed genes, ~40% of which carry canonical NMD-inducing features, implicating pathways linked to calcium signaling, phagosome formation, and glial development. A rescue experiment using the CalEx calcium extrusion pump demonstrates partial restoration of synaptic strength and anxiety behavior when astrocytic calcium is normalized.

      The study addresses an important gap in our understanding of RNA regulation in glial cells, and the overall conceptual framework is well described. The experimental design is generally appropriate, and the multi-pronged approach lends the main claims a degree of validity.

      Strengths:

      (1) Novelty: This is the first study to systematically examine NMD function in astrocytes in vivo. The identification of astrocytic NMD targets via RNA-seq combined with an NMD-inducing feature classifier is a meaningful methodological contribution.

      (2) Multi-method approach: The authors combine morphological analysis (Imaris 3D reconstruction), synaptic markers (PSD-95, LAMP2 engulfment assay), spine density measurements, acute slice electrophysiology, two-photon calcium imaging, behavioral testing, and transcriptomics. The convergence across these methods strengthens confidence in the claims.

      Weaknesses:

      (1) While the transcriptomic analysis is a valuable addition, the connection between specific NMD targets and the observed calcium phenotype remains largely correlational. The authors identify Gabbr2 and Adora1 as upregulated candidates with canonical NMD features and speculate that their elevated expression drives aberrant calcium signaling. However, no validation (e.g., qRT-PCR or protein-level confirmation) of these candidates is presented. The mechanistic pathway between NMD disruption and elevated calcium is thus inferred from pathway analysis rather than demonstrated. This is a significant gap between the transcriptomic and physiological arms of the study, and the authors should be more explicit about this limitation or, ideally, provide at least one validated target.

      (2) The reduction in astrocyte surface area in cKO mice is interpreted as contributing to reduced synapse contact and engulfment capacity. This is a reasonable hypothesis, but the study does not directly demonstrate that reduced astrocyte territory correlates with reduced synaptic coverage at the level of individual cells or brain regions. The temporal sequence of these events is unknown. Do morphological deficits precede synaptic changes? Clarification and qualification of this causal chain in the Discussion would strengthen the manuscript.

      (3) LFS-induced LTD is unaffected, while mGluR-LTD is reduced. This is intriguing and potentially informative about astrocyte contributions to distinct LTD mechanisms, but the difference receives limited discussion. Given the relevance of mGluR signaling to calcium dynamics and the identified pathway enrichments (GPCR signaling), this specificity deserves more attention.

      (4) The CTRL + CalEx condition is included in the EPM experiment but not in the electrophysiology or calcium imaging experiments, making it difficult to fully assess whether CalEx itself has off-target effects on synaptic transmission or anxiety in wild-type animals. The CTRL + CalEx EPM data (Figure 7F) appears to show a modest reduction in open arm time relative to CTRL, which, if robust, would suggest that excessive calcium reduction in astrocytes is also anxiogenic. This finding would be physiologically relevant and deserves comment.

    3. Reviewer #2 (Public review):

      Astrocytes are highly responsive to their environment and play a range of critical roles in brain function. Lituma et al. theorize that one mediator of that responsiveness is the regulation of RNA stability. They therefore undertake an assessment of astrocytes missing Upf2, a protein required for mRNA degradation via nonsense-mediated decay. This is an interesting study, approaching astrocyte biology from a novel angle. The authors take on an ambitious set of experiments, spanning morphological assessment, synaptic engulfment, electrophysiology, behavior, and calcium imaging.

      The authors show convincing data that knocking out Upf2 in astrocytes impairs synaptic plasticity, affects behavior, and changes the complement of astrocytic mRNA. These results, in and of themselves, are intriguing and suggest that NMD is an important biological process in astrocytes, warranting further study.

      My primary concern is whether the authors may be largely studying dying cells. The idea that NMD disruption has a dramatic effect on astrocyte morphology is an intriguing idea, but it is not fully established here. The nuclei in the example cKO morphology images appear small and/or fragmented. This raises concerns that the authors did not ensure that they had the full 3D morphology of the astrocyte in the section, and the cell is in part cut off, which would compromise any data on the morphology. The authors state that the tissue was sectioned at 70 um. The diameter of an astrocyte in the adult mouse brain is typically between 50 and 70 um. Unless astrocytes are perfectly positioned in the center of the slice, at this thickness, the majority of astrocytes will almost certainly be partially cut off. More detail on how cells were chosen and what quality control metrics were implemented would alleviate concerns here. An alternative possible explanation for these small/fragmented nuclei is that cKO astrocytes may be unhealthy to the point that they are actively dying. Using the transgenic ZsGreen label, the authors state that they observe a size change (Figure S4); this is not readily apparent and is not quantified in any way. It does appear from these images that there may be a loss of some astrocytes; cell death, which would also be an interesting finding, is a fundamentally different process than morphologic restructuring in living cells. The authors do attempt to count astrocytes (Figure S6B), but do so with GFAP. This is a fundamentally flawed approach. Because GFAP is not readily detectable in most healthy astrocytes in most gray matter regions, GFAP should not be used to quantify astrocyte numbers; this experiment should be repeated with a better marker, such as Aldh1l1, Sox9, etc.

      Synaptic engulfment: This is an extraordinarily high degree of engulfment in the control animals compared to many published studies, leading to concern as to the technical approach. Indeed, the overall low level of PSD-95 signal in control conditions in adult mice is concerning as to the technical accuracy of the approach. It is unclear exactly how the investigators labeled the astrocytes; presumably via the ZsGreen label, but it is never stated, and the only images shown are the highly processed Imaris renderings. The small astrocytic processes, or leaflets, that make up the vast majority of the astrocytic arbor are on the order of 100nm in diameter. The processes shown in Figure 2B are, according to the scale bar, at least 20x that size. It is difficult to have much faith in these results as currently presented.

      The signal-to-noise ratio of the GCaMP experiments is worryingly low, likely responsible for the abnormally low dF/F in all conditions and the lack of significant change between control and CalEx, when control astrocytes should show a much higher GCaMP signal than any CalEx-expressing astrocyte. That said, the higher Ca++ in Upf2 KO astrocytes is intriguing. Given the roles of elevated calcium in cell death, this may reflect cells that are unhealthy to the point that they are starting to die.

      The authors conduct a FACS-based analysis of astrocytic mRNA from control vs Upf2-KO, with intriguing results. An important caveat, though, is that a large amount of astrocytic mRNA is in the processes. If mRNA stability is being actively and rapidly regulated, it seems likely that the mRNA in the processes would be the most relevant population of regulated mRNA. FACS-based approaches to astrocyte purification will, as robustly shown elsewhere, strip off those processes. Particularly given that the authors have shown that the processes may be the most actively changing astrocytic compartment with Upf2 KO, this is a strange choice of technique vs. something like Ribotag that would preserve the mRNA in processes. At least, there should be some discussion regarding using FACS for this analysis and the consequences for profiling mRNA in astrocytic processes.

      Minor points:

      (1) The use of the Aldh1l1-CreER mouse is a strong choice and has been shown to be highly astrocyte-specific. Combining that transgenic mouse with viruses driven by different forms of the GFAP promoter is quite bizarre in several ways. First, GFAP-dependent AAVs have been shown repeatedly to have significant neuronal leak. Second, these mice are, in all cases, receiving two different viruses, driven by different forms of the GFAP promoter, and the non-Cre virus is not Cre-dependent (vs. a much more standard approach of using a Cre-dependent second virus to ensure that all analyzed cells received both viruses). The authors mention that "this experimental design ensures that phenotypes are not caused by an acute effect of tamoxifen." It is certainly true that tamoxifen is not a biologically neutral molecule. However, the mice still receive tamoxifen, both in these morphology virus experiments and in almost all other experiments. This experimental approach is not inherently bad, nor does it necessarily invalidate the data (although the near-certain neuronal contamination due to the GFAP promoter-driven viruses is a concern). It is, however, convoluted in ways that appear unnecessary. If there is a strong rationale for this approach beyond the tepid explanation already present, it should be explicitly mentioned.

      (2) The characterization of the knockout is incomplete. While the authors should be applauded for their attempts to phenotype the cells in which they observe Cre-mediated recombination, there are issues with their technical approach. Most importantly, and an issue that affects other analyses in the paper as well: the vast majority of astrocytes in the healthy cortex do not express GFAP. Therefore, using GFAP to claim high astrocyte specificity and efficiency is a fundamentally flawed approach. Second, MBP is a myelin marker, not a cytoplasmic marker, and would not successfully colocalize with a cytoplasmic marker like ZsGreen even if recombination in oligodendrocytes did occur. Third, recombination at one set of LoxP sites is not a reliable indicator of recombination at other sites. Recombination efficiency is highly dependent on the spacing between the LoxP sites and cannot be reliably extrapolated to other floxed genes without validation. Finally, the most likely culprit for off-target recombination with Aldh1l1-CreERT2 (or other astrocyte-selective Cres, and certainly the GFAP-based viral promoters) is neurons, which the investigators did not test for. Neuronal Aldh1l1-CreERT2 leak is most likely to occur in the hippocampus. With the images shown in Fig S3, it is unclear whether it is possible to convincingly colocalize Upf2 staining with a cytosolic marker of all astrocytes, such as Aldh1l1 or S100b, but such data would be more appropriate. An alternative approach to validation would be in situ hybridization.

      (3) Supplementary Table 2 should include gene IDs, not just Ensemble IDs.

      (4) It is not fully clear what the investigators are denoting as a spine in Figure 2E; the two images do not appear to have the large degree of difference that the quantification suggests. The oversaturation of the signal complicates assessment.

      (4) A more detailed discussion of the rationale behind the timeline would be helpful. What is the half-life of Upf2, and how rapidly do NMD genes build up upon Upf2 disruption? In particular, in the case of virus experiments, the timeline is quite fast: ~2.5 weeks from injection to analysis. ssAAV expression takes over a week to reach appreciable levels.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigate mRNA targets of the nonsense-mediated decay (NMD) pathway in astrocytes and link the dysfunction of NMD in astrocytes to aberrant synaptic transmission that has downstream effects on behavior. Specifically, they find a link between the aberrant synaptic transmission with elevated spontaneous calcium signaling in astrocytes, and functionally they demonstrate that manipulating astrocyte calcium signaling with CalEx modulates astrocyte calcium signaling towards wildtype levels and improves anxiety behavior. They investigate the astrocyte calcium signaling changes in Upf2 conditional knockout mice in several brain regions that have been linked to anxiety behavior, including the hippocampus and prefrontal cortex. They also observe aberrant astrocyte calcium signaling in the visual cortex, demonstrating that dysfunction of the NMD pathway in astrocytes has widespread effects on synaptic transmission in various brain regions. This work identifies, through RNA-Sequencing, potential mRNA targets of NMD in astrocytes, and shows that pathway enrichment of these targets highlights calcium signaling. Altogether, this work highlights the importance of the basic cellular process of NMD in astrocytes, which are known to have extensive local translation of proteins in their perisynaptic processes. NMD may be particularly important in astrocytes due to their intimate association of processes with neuronal synapses, and the authors suggest that alterations to NMD function in astrocytes may be an important avenue for future investigation in neurodevelopmental disorders.

      Strengths:

      Altogether, this work is a critical foundation for future research into astrocyte contributions to neurodevelopmental disorders. The authors do a thorough characterization of astrocyte conditional Upf2 knockout mice in several brain regions. They present a complete story that connects molecular events (NMD pathway regulation of mRNA degradation) to astrocyte regulation of circuit activity to organismal behavior. The electrophysiological analysis is thorough, and the manipulation of calcium activity ties astrocyte calcium activity to anxiety behavior. The RNA-sequencing dataset is useful to the scientific community and provides a resource of candidate molecules that might be dysregulated in neurodevelopmental disorders.

      Weaknesses:

      The study suffers from some overstated claims and a lack of statistical rigor in some experiments, as detailed below.

      (1) The title states that "Astrocytic Nonsense-mediated mRNA decay regulates calcium signaling to support synapse function and restrain anxiety". The term "restrain anxiety" implies that the NMD pathway has a direct effect on a molecular switch to control anxiety. Anxiety behavior is a complicated process, controlled by many biological phenomena and synaptic transmission in the circuit as a whole, and is not directly linked to a specific NMD mRNA target. This title is overstating the findings of the study.

      (2) In general, the first figures (1-2) suffer from low power (N = 3) and statistical rigor. The statistics are inflated by analyzing individual fields of view and per-cell data rather than performing the statistics on the average of biological replicates. It is preferable to show the biological replicate data so that readers can observe the natural biological variability between replicates.

      (3) The claim that astrocytes have decreased engulfment of synapses in the Upf2 conditional knockout mice is not strongly substantiated by the data. The resolution of confocal microscopy and the static nature of histological images make it difficult to measure synaptic engulfment as an active process. Additionally, the metric of quantifying the % occupancy of PSD95 puncta within the total astrocyte volume may be skewed due to overall differences in cell size (shown in Figure 1). There is not much discussion of how a decrease in astrocyte engulfment of synapses may lead to decreased synapse number. To the contrary, one might expect decreased engulfment to result in increased synapse density.

      (4) The authors use Gfap as a marker to count astrocyte cell number and assess if there are changes in cell number between genotypes (Figure S6). However, Gfap does not label all astrocytes in the cortex and, in fact, is rather an aberrantly expressed marker in conditions of inflammation, as opposed to the hippocampus, where Gfap is basally expressed in all astrocytes. In the cortex, there seems to be a trend for reduced Gfap in the conditional knockout mice, which may suggest differences in astrocyte molecular signatures rather than cell numbers. Another astrocyte marker, like Aldh1L1, will be more accurate to assess this question histologically.

      (5) The authors state that "Preventing abnormally high basal calcium activity in NMD-deficient astrocytes restores normal excitatory synapse function...". However, this claim is not substantiated by the data. CalEx manipulation certainly shifts the input-output curve but does not restore to wildtype baseline levels (Figure 6E). Additionally, synapse number does not appear to be restored to wildtype levels (Figure 6D - although the p-value for this comparison is now shown). The investigators do observe improvements in anxiety phenotypes, suggesting there is some modulation of circuit activity, but the claim that CalEx manipulation restores baseline synaptic transmission is not supported.

    5. Author response:

      We thank the reviewers for their careful reading of our manuscript and for providing positive, constructive feedback. In particular, we thank the reviewers highlighting the several strengths of our study.

      To address the reviewers’ major concerns, we will revise the presentation of our main findings (specifically data/animal vs data/ROI), provide more clarity in the Results, Methods, and Discussion sections, and modify the title to better reflect these nuances.

      Additionally, we will perform the following new experiments:

      (1) Astrocytic Marker Validation: To further confirm comparable astrocyte cell counts between the CTRL and Upf2-cKO conditions, we will perform immunostainings using Aldh1L1, Sox9, or S100b instead of GFAP.

      (2) NMD Candidate Validation: To validate top candidate NMD target transcripts, we will perform immunostainings or qRT-PCR for Gabbr2, Adora1, S100b, or Cldn9.

      (3) Sample Size Expansion: To strengthen the morphological and PSD-95 quantifications, we will increase the sample size (N) by incorporating additional animals.

      (4) Mechanistic Timeline & Phenotype Linkage: We value the reviewer’s comment regarding the timeline of morphological and Ca<sup>2+</sup> phenotypes. To gai insight into whether these phenotypes are independent or linked, we will perform 3D reconstructions in CalEx conditions to assess whether Ca<sup>2+</sup> restoration rescues astrocyte morphology in Upf2-cKO mice. This will allow us to determine if increased Ca<sup>2+</sup> activity is upstream of the morphological alterations. Taken together, we believe that incorporating these manuscript revisions will strengthen the clarity and conclusions of our work. We thank the reviewers for their time and careful evaluation of our study.

    1. eLife Assessment

      This is a valuable paper looking at nanoscale organization of the membrane associated periodic cytoskeleton in mouse sciatic nerve axons. Despite previous studies, the precise organisation of the structure remains unclear, especially in vivo, and this manuscript significantly adds to this knowledge with solid data. An unexpected observation is the presence of discrete nanoscale clusters, regularly distributed around sections of axons. However, the paper misses a description of these clusters along the longitudinal axis of the axons.

    2. Reviewer #1 (Public review):

      Summary:

      The article "Nanoscale organization of beta-II spectrin within segments of the membrane-associated periodic skeleton in mouse sciatic nerve axons" by Gazal et al. looks into the organization of the spectrin scaffold in mouse sciatic nerves using super-resolution microscopy. It is now well established that axons, across species, contain a membrane-associated periodic scaffold mainly composed of circumferential actin filaments and longitudinally arranged spectrin tetramers. While super-resolution imaging of neurons in cell culture is relatively easy, exploring the ultrastructure of myelinated axons in intact nerve fibers is a daunting task. Nevertheless, the authors have attempted this by fixing and preparing cross-sections of sciatic nerves. They have then tried to quantify the fluorescence intensity patterns of specific components, especially that of labeled beta-II spectrin and have analysed its distribution.

      One of the main findings is that spectrin is distributed along the axonal periphery and along the outer part of the myelin sheath. By labelling multiple cellular components and using intensity analysis, the authors show the sequence of structural organization of a few key components. They see that, unlike in the case of axons in culture, the axonal cross-sections within the sciatic nerve deviate significantly from a circular shape. They then use 3D-dSTORM to investigate the distribution of beta-II spectrin along the axonal circumference. They see that this distribution is very heterogeneous, both in the sizes of spectrin puncta and their arrangement along the periphery. The amount of spectrin scales linearly with axonal circumference.

      Strengths:

      Super-resolution imaging of axons of intact nerve fibers to investigate the organization of beta-II spectrin.

      Weaknesses:

      While most of the findings, like the spatial distribution of spectrin and related components, are reasonably well supported by data, I have concerns regarding the subsequent claims made in the article. The detection of axial periodicity based on the observation of a peak in the inter-tetramer spacing distribution is not very convincing, and a 3D representation (or a video of 3D reconstruction) would have been better. And so are the claims on characteristic spectrin spacing of 200 nm along the axonal circumference. A peak in the distribution does not imply a periodic arrangement.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting paper by the Unsain lab looking at the nanoscale organization of the membrane-associated periodic cytoskeleton in mouse sciatic nerve axons. The precise organization of the structure remains unclear, especially in vivo, and this manuscript significantly adds to our knowledge of this important structure. While some of the findings in the study are somewhat expected (though still valuable to see in an in vivo setting), an interesting observation is the presence of discrete nanoscale clusters that scale up with the size of the axon, which challenges previous assumptions.

      Strengths:

      Strong, convincing data; clever combination of imaging and analytical tools to make novel points; well written; excellent composition of figures.

      Weaknesses:

      (1) Figure 2A/3A: The large and small clusters of spectrin, as seen in cross sections, are unexpected and novel. The authors have done a clever job of combining imaging and analyses, but some things are still unclear. First, the authors should be consistent in their language when they talk about the spectrin clusters. Recommend precise language to define the small and large clusters when they first appear in the text, and then use the definitions consistently throughout the text. Second, based on the data shown, one does not get a clear idea of how the small and large clusters are organized along the longitudinal axis of the axon. In that context, are Figure 2B and C from imaging along the longitudinal axis? If not, it's unclear how the authors can conclude that the spectrin assemblies have a distance of ~170 nm along the linear axis. In general, a perceived limitation of this study is that while the authors have done a good job looking at cross sections, there is no information on the longitudinal distribution of spectrin in these axons. Looking at both cross- and longitudinal sections would also clarify details about the large spectrin clusters. For instance, are they small sausage-like structures, or long rods of spectrin running along the length of the axon? One assumes that all the analyses in Figures 3 and 4 are from the small clusters. Can the authors do a similar analyses of the large clusters? Finally, a schematic model showing both cross- and longitudinal- sections would make things clearer, but the authors would need to show the longitudinal data for that.

      (2) It is interesting to think that the larger spectrin accumulations may be similar to the condensate-like structures seen by Boyer et al., as the authors mention in the discussion. In that context, it is possible that these focal accumulations are local reservoirs of spectrin that are also seen in mature axons (indeed, these accumulations were also seen in mature axons in the Boyer et al. paper, and they also speculated that these accumulations may be local reservoirs). Can the authors check if actin/adducin is also present in these larger spectrin accumulations?

      (3) While talking about the nanoscale clusters, it is important to specify that the authors are talking about circumferential clusters. Though the writing is excellent, one still does not get the precise definition of "clusters" from just reading the abstract, and it would be good if the authors could work on that more (I recognize that this is not easy to do).

    4. Reviewer #3 (Public review):

      Summary:

      In the presented work, the authors investigate spectral staining in axons of the sciatic nerve, where the MPS has been detected before using STED microscopy. They employ 3D-dSTORM in tissue sections and analyze the data, measuring localization of clusters on the axon perimeter and the relative distribution of those. From these data the conclude that large gaps in spectrum localizations exist and that clusters around the axon exist that are spaced at 200nm.

      Major Comments:

      (1) The presented data are at times overinterpreted, and the discussion lacks a critical view of the data. For example, the statement "...Unlike previous suggestions from qualitative evidence in cultured neurons (REfs), βII‑spectrin distribution in MPS segments of peripheral nerves is discontinuous, with extensive stretches of the perimeter lacking βII‑spectrin." is quite strong, given it is based on immunofluorescence staining and dSTORM microscopy in tissue. Absence of evidence of staining is not evidence of absence.

      (2) The authors claim in the abstract that "The number of these clusters scales linearly with the axonal perimeter, maintaining a constant membrane occupancy of ~20% across varying axon diameters." Again, this is from a cut through an axon, while measuring the density of clusters on the perimeter. If they claim area occupancy, an area should be imaged, and the dots (clusters) should be measured in surface coverage in a 2D projection of the axonal surface.

      (3) In general, this reviewer suggests being a bit more moderate in statements such as: "These findings challenge simplified models of the MPS based on cultured systems and demonstrate that the MPS in peripheral nerves is composed of discrete structural units." These statements are bold from the relatively few measurements in a single method and a single viewpoint. Especially when considering that techniques such as dSTORM depend extremely highly on labeling density, and apparent clustering of localization is highly prone to misinterpretation. If the authors desire to make such statements, working with endogenously labeled protein would be warranted. The authors should at least hedge such statements.

      (4) If the authors want to make statements about general organization, why do they not compare adjacent cuts through the axon? If there are continuous spectrin filaments, the clusters should appear at the same site across repeated cuts through the axon.

      Besides this, this reviewer welcomes the effort that has been made to establish dSTORM in tissue sections and to investigate the MPS in native tissue.

    5. Author response:

      We sincerely thank the editors and reviewers for their overall positive assessment and constructive feedback on our manuscript detailing the nanoscale organisation of βII-spectrin of the membrane-associated periodic skeleton (MPS) in mouse sciatic nerve axons. Their perspective and comments will help refining the manuscript.

      A common comment by the reviewers relates to the description of the characteristic longitudinal periodicity of the MPS. We value these comments, which we believe are motivated by the fact that the longitudinal periodicity of the MPS is undoubtedly the most studied and prominent feature of the MPS in cultured neurons. However, the main goal of the present project was to describe how βII-spectrin is organised in the transverse axis of individual segments of the MPS in nerve tissue. This is why we utilised cross-sections of the sciatic nerve, hence achieving the best resolution possible in that plane, at the expense of the resolution in the axial axis. Furthermore, this study clearly shows that the transverse morphology of axons, and thus of the MPS, of neurons in the tissue is highly irregular, in comparison to cultured neurons. This imposes an extra challenge to observe correlated longitudinal structures when the observation length is limited, as in our studies. Nonetheless, to improve this aspect of the manuscript, we will revise our data and previous evidence, clarify the methodological trade-offs made, and make our interpretations more accurate.

      Additionally, we will clarify several imaging- and definition-related inquiries, including tests for insufficient staining, the interpretation of βIII-tubulin staining, the assessment of axon–glia boundaries, and consistency in the use of terms like ‘clusters’ and ‘periodicity’, among others.

      We believe these and other revisions will substantially strengthen the manuscript and comprehensively address the reviewers' feedback.

    1. eLife Assessment

      This valuable study shows the impact of the metabolic state of bacteria on phage infection. The experimental results, based on various phages infecting E. coli, are convincing and consistent with a two-step adsorption mathematical model. This study should be of interest to the communities working on cell metabolism and on host-pathogen interactions.

    2. Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. The mechanistic interpretation offered by the model highlights an interesting trade-off between adsorbing to a metabolically active host and discriminating between active and inactive hosts that, somehow, a phage has to optimize. It would be interesting, in the future, to investigate how different phages optimize this task.

      Comments on revised version.

      I am happy with how the authors have addressed all the comments. The manuscript is much clearer and more readable and the previous overstated claims have been removed/clarified.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phage that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well designed and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication and the paper shows that host physiology has to be taken into account at all steps of phage infection.

    4. Reviewer #3 (Public review):

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, 𝜙80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      As I have stated in the first submission of this manuscript, the data presented by the authors clearly supported the principal conclusion of the study. The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully use a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Comments on revised version.

      In this revised version, the authors have successfully resolved all of my comments. I appreciate that the main text has been majorly revamped, which greatly helps the readers follow the motivation behind the experiment and analyses, and interpret the data. I agree that the revised terminology choice "commitment to infection", instead of the previous interchangeably used "adsorption"/"entry", is much more logical, considering the experimental data. I also commend the authors for writing the modeling part in a very clear, pedagogical, and instructive manner. Overall, I believe that this manuscript will be valuable to those who are interested in phage-bacterial interactions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource-limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. While the model is sensible from a mechanistic perspective, the experimental evidence that supports how each model's rate is affected by the cell metabolic state is weak, as only ratios of these rates can be inferred from the data.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phages that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well-designed, and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication, and the paper shows that host physiology has to be taken into account at all steps of phage infection.

      Weaknesses:

      There are some concerns about the experimental setup and which conclusions can be drawn from it:

      Before phage infection, bacterial cultures are grown to exponential growth, washed, and then resuspended with glucose or arsenate-azide for 10min. It is however, questionable that 10 minutes is enough to simulate high and low metabolic states realistically. 10 minutes seems to be quite short to go from exponential growth to a low metabolic state, given the transcriptional memory of previous environments. It seems more likely that the population will be quite heterogeneous, with cells in various states of transition towards low metabolic states.

      While we agree with the reviewer that during metabolic transitions there may be a period in which the population is heterogeneous, with cells in different stages of transition toward a low metabolic state, the 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyper diffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027). We have also corrected the DOI for this reference in the manuscript. Furthermore, the ATP pool of log-phase E. coli turns over several times per second (Holms et al., Arch. Mikrobiol. 1972, http://dx.doi.org/10.1007/BF00425016). We therefore assumed the bacteria were energy depleted after 10 minutes. We have clarified this point in the revised manuscript.

      Given that arsenate and azide inhibit cellular metabolism, i.e., have antimicrobial effects, cells might not just downregulate metabolism but also activate the stress response, and this causes some of the observed effects on phage adsorption. Therefore, the 'low metabolic state' of the cells in this paper could mean that cells are starved or that they are stressed or both.

      The reviewer is correct. We don’t exclude indirect effects. However, as nutrients were removed from the bacteria by washing and energy metabolism was inhibited by the addition of arsenate and azide, we assumed a stress response requiring biosynthesis would be unlikely to occur.

      The abundance of receptors could change between the high and low metabolic media conditions and contribute to the observed differences in adsorption, while the authors seem to assume in their model that the initial adsorption rate always remains the same.

      We do not think that the observed differences in adsorption are explained by a change in receptor abundance. In a previous study using the same experimental protocol as in the present work, phage λ was compared to the metabolically insensitive mutant λh (Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119). If the lower adsorption in the low-metabolic condition were caused by a reduced number of receptors, then λh should also have shown a lower adsorption rate under the same condition. Instead, no measurable effect on λh adsorption rate was observed. We therefore conclude that the effect is not explained by changes in receptor number on the timescale of the experiment. We have clarified this point in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, ɸ80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy-depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      Strengths:

      The data presented by the authors clearly supported the principal conclusion of the study ("Viral commitment to infection depends on host metabolism"). The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully used a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Weaknesses:

      (1) The authors isolated and measured the numbers of free phages in the medium after infection of bacteria under different treatments. These measurements were analyzed in two different ways: (1) simply as ratios (corrected/normalized using different controls), and (2) fitted using a simple mathematical model. I have concerns regarding both analyses.

      (1.1) For the first method, having different time points at which the sample of each phage is collected critically complicates data interpretation. As one incubates the phage-bacteria mixture for a longer time, more infection occurs, and the number of phages collected from the mixture decreases. Therefore, the different incubation time forfeits the goal of "a systematic and quantitative comparison across different phages [...]", just as the authors self-criticized. Conceivably, the authors could have used the shortest measurement time for all phages (i.e., 10 minutes, as for phage λ). Alternatively, the authors could have applied a systematic criterion such as half (or any other fraction) of the latent period of each phage, which would still "maximize the incubation period while ensuring that manipulations were completed before the first infection cycle concluded". In my view, the seemingly arbitrary measurement time for each phage renders the entire first analysis very challenging to interpret. It also goes against the author's proposition that the protocol was "standardized" or "consistent". It is not clear what the readers are supposed to take away from this first analysis, or rather, which evidence, finding, or conclusion the manuscript would lose if the authors only presented the modeling-based analysis.

      (1.2) The second method of analysis sought to remove the dependence of the measurements on time. I completely agree with this goal, and the findings extracted from this analysis significantly contributed to the merits of this manuscript. However, the authors achieved this goal using a single time point for each phage to calculate the infection rate (η). As shown in Figure S3, each of the phage depletion curves is anchored by only one data point (note that the P(t)/P(0) = 1 at t = 0 is assumed, not measured). This goes against the typical way this collision model is used in the literature, where a time series is measured and used to fit the model (e.g., DOI 10.1007/978-1-60327-164-6 18, or more recently, PMID 39700139). This practice in the current manuscript reduced the robustness of the inferred η values. This problem is exacerbated by assumptions used by the authors in formulating this model. For instance, the authors used a constant value for the bacterial concentration, B, because "bacterial growth and lysis were negligible" (lines 135-136). However, considering that the bacteria were cultured at 37oC in a very rich medium (first in YT broth, then in 2% glucose), the measurement times of 20, 30, and 55 minutes are most likely one or a few generations of bacterial growth and division.

      Related note: I suggest that one of the panels in Figure S3 should be moved to the main text, since it is critical to the second method of analysis.

      We would like to clarify that the manuscript does not present two separate methods, but rather one method presented in two steps: a first step with results that are directly tied to the experimental measurements and show whether the effect is present for each phage, followed by a second, analytical step that makes the results comparable across phages.

      The first step presents the ratios because they directly reflect the measurements performed in the experiment and allow the reader to see the effect of the metabolic state for each phage in contrast to its control. We agree that these ratios are time-dependent and therefore not suitable for quantitative comparison between phages. Their purpose is to illustrate the experimental outcome and to show that the effect is present (or absent) on a per-phage basis not to compare magnitudes across phages.

      We then follow this with the second step, allowing the reader to follow the logic of the analysis. The analytical step that follows does not represent a second method, but a continuation of the same analysis. Here, we remove the time-dependence specifically in order to make comparison of the effect across phages possible, by connecting our results to standard measures such as the adsorption rate η. Importantly, P(0) is measured for every phage in every experiment. The only modeling assumption used (a standard one in the field) is the exponential form for the decay in free phage number, which naturally yields P(t)/P(0) = 1 at t = 0.

      Regarding the reviewer’s concern that bacterial growth may not have been negligible over the relevant time window, we note that recent work on rich-to-minimal growth lags in E. coli reports substantial delays before growth resumes after nutrient downshift. One 2023 study (Wu et al., Nature Microbiology 2023, https://doi.org/10.1038/s41564-022-01310-w) considering wild-type E. coli shows in Fig. 2c a lag of up to about 2 hours after a shift from MOPS minimal medium with 0.2% glucose plus 18 amino acids to the same medium without amino acids. Another 2023 study (Zhu and Dai, Nature Communications 2023, https://doi.org/10.1038/s41467-023-36254-0) examining both rel+ and rel− strains reports a growth lag of about 49 minutes for rel+ and more than 5 hours for the relA deletion strain. While these conditions are not identical to ours, they support the general point that growth does not immediately resume after such shifts. We therefore think it is unlikely that, following transfer from YT, the cells underwent one or a few full generations during the time window of our adsorption measurements.

      On the related note: Following the comments of all reviewers on Figure S3, we have decided to remove it to avoid confusion.

      (2) The data were able to distinguish phages that successfully infected bacteria and those that remained free in the medium, and the authors appropriately interpreted the data as such throughout the Results section. However, in the Discussion (starting from the very first sentence, line 172), the authors used terms that include "adsorption" and "entry" more interchangeably (for example, see the three sentences in lines 310-313, for "viral entry efficiency is shaped by [...]", then "adsorption kinetics modeling"). I do not see how the authors' data could distinguish between adsorption (the phage particles attaching to the outside of the cell) and entry (the phage DNA being injected into the cell). Conceivably, any phage particles that irreversibly attach to a cell but do not yet inject their genome into the cell would still be removed from the medium and therefore not quantified. Another example: in lines 189-191, the authors interpreted that "[...] when the bacterium is in a low metabolic state, the phage does not bind irreversibly to the host", but how do the authors eliminate the case of no phage binding (i.e., the reversible step) to begin with?

      We agree with the reviewer that our use of the terms adsorption, entry, and infection should have been more careful. Our experiment can only identify the irreversible commitment of phage to a host cell. We have therefore revised the text to refer consistently to phage commitment.

      Similarly, in lines 283-293, how do the authors delineate whether energy depletion would increase the k_off term or decrease the k_inj term, because either would result in more free phages in the medium as observed in the data? I believe that the writing of the Discussion, as it stands now, is doing a disservice to the conclusions presented in the Results section.

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: if energy depletion leads to reduced commitment, this can arise either because k_off increases, because k_inj decreases, or because both change, as long as k_off/k_inj becomes larger in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also leads to the trade-off now discussed in the revised manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become significantly larger in the inactive case. Depending on whether this is achieved through changes in k_off or k_inj, the cost of discrimination appears either as slower commitment or as additional energy dissipation. We agree that the previous wording overstated the mechanistic interpretation, and we have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (3) The authors presented an argument that performing infection of all five phages in the same condition is an advantage, allowing for comparison across different phages. While this goal is a completely valid one, it is difficult to reconcile that with the fact that different phages require different optimal conditions for successful infection. For instance, phage T5 famously requires Ca2+ for successful infection into the host bacterium (and later successful replication); see PMID 13174489. However, all infections were performed in TMG, which lacks Ca2+. Perhaps the absence of T5 dependence on the host metabolism is because the infection condition used by the authors was not optimal for T5 to begin with? Similar arguments could be made for other phages.

      Our study alone cannot eliminate that possibility. However, we have cited multiple previous studies, for example references citing Braun et al., showing that T5 remains insensitive to the host metabolic state under different buffer conditions. We therefore believe it is unlikely that the lack of metabolic dependence we observe for T5 is simply due to suboptimal infection conditions.

      (4) Whereas the manuscript examined five coliphages, only phage T5 and phage λ were discussed extensively. I believe some discussion points for these two phages need clarification.

      We focused our discussion on the phages T5, λ and φ80 because these are the phages for which similar effects have been reported previously in the literature. This allowed us to connect our findings directly to existing work and to discuss mechanistic hypotheses in a meaningful comparative framework. For the remaining phages, to our knowledge no prior studies have examined their behavior under comparable metabolic conditions, and therefore a similarly detailed discussion would have been speculative. Nevertheless, all five phages are treated equally in the presentation of the experimental results and in the quantitative comparison of adsorption rates.

      (4.1) Phage T5: The data obtained by the authors show that the infection rate of phage T5 is not impacted by the metabolic state of the host cell. Considering that the authors used the terms "infection", "adsorption", and "entry" interchangeably to refer to the irreversible commitment of a phage to a host cell (see point 2), this discussion regarding phage T5 lacks one critical literature context: DNA entry of phage T5 is known to occur in two phases (first-step transfer and second-step transfer). Critically, the second step can only occur if phage proteins encoded by the phage DNA transferred in the first step are expressed (see PMID 10577483 and the cited papers therein). In that context, metabolic poisoning of the host bacteria should have impeded T5 infection. The authors should comment on this point.

      As the reviewer pointed out, our usage of the terms infection, adsorption, and entry should have been more careful. Our experiment can only identify irreversible commitment of phage to a host cell. For T5, we expect that this irreversible commitment already occurs upon first-step transfer of phage DNA. As a result, even if second-step transfer is impeded under metabolic poisoning, our method would not resolve that effect. We have added this clarification to the revised manuscript.

      (4.2) Phage λ: The experiment using phage λ in this current study shares many resemblances to that in Brown et al. 2022. That feature alone is not a problem, but at many places in the text, the writing is ambiguous as to whether it is discussing the results in Brown et al. 2022 or in the current manuscript. I am giving three examples below, but this is not exhaustive: (i) Lines 67-69, there is no Brown et al. 2022 reference immediately after "a mutant phage variant (λh) could bypass this dependency [...]" (not just in the previous sentence); (ii) Line 228 should clearly say "Our previous findings suggested that phage λ is capable of [...]", since it concerns Brown et al., 2022, not the current study; and (iii) Lines 245-246, there is no Brown et al., 2022 reference immediately after "we observed that a mutant variant [...] even energy-depleted host" (without a reference, it reads like the authors "observed" that finding in this current manuscript).

      The reviewer is right. In those places, the text was ambiguous as to whether it referred to the present study or to Brown et al. (2022). We have now inserted the reference at the relevant points and revised the wording where needed to make this distinction explicit.

      Also, regarding phage λ: The discussion between line 230 and line 249 is very interesting, but since it concerns the differences between λ PaPa and Ur-λ, the authors should consider mentioning and discussing a very relevant recent study, PMCID: PMC6312755.

      We agree that the study by Guan et al. is very relevant and interesting. However, our point in this part of the Discussion is only to clarify that we used λ PaPa and not the originally isolated λ strain. We have therefore limited the discussion here to that distinction.

      (5) Control experiments, or references to prior studies, are needed to support that the As/Az treatment at this concentration and duration (at least 10 minutes) is sufficient to deplete the metabolic state of the cell. For instance, this can be shown by impeded or null cell growth, arrested motility (using a standard swimming assay), or a fluorescent reporter for the energetic state of the cell.

      The 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyperdiffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027) where the effect was assessed by monitoring the rate of movement of the λ receptor on the bacterial surface. We have clarified this point in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As mentioned earlier, I found the paper interesting and addressed an important and significant knowledge gap.

      My biggest concern is about the interpretation of the experimental data in light of the two-step model. In particular, around line 286, it is stated "k_inj is more sensitive to metabolic state than k_off". Assuming k does not depend on metabolic state, which is a fair assumption, the equation for eta only depends on the ratio between k_inj and k_off and not on the individual parameters separately. Consequently, there is no way of saying which one of the two is more affected by metabolic state, unless the model already assumes that k_off is not influenced by metabolic state. The results could equally be explained by k_inj decreasing in metabolically depleted cells, or k_off increasing in such cells. If this is an assumption of the model, this should be clearly stated and not reported as a consequence of the data, as it is at the moment. Also, how does this mathematical model connect to the fitting function used in Figure 2b?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: discrimination requires that k_off/k_inj be larger for inactive hosts than for active hosts, such that commitment is specifically reduced in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This introduces a trade-off: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination requires this ratio to become significantly larger for inactive hosts. If this is achieved through changes in k_off, discrimination comes at the cost of slower commitment by allowing more time to leave; if it is achieved through changes in k_inj, it can preserve fast commitment to active hosts but requires additional energy dissipation in order to actively modulate commitment. We have therefore revised the text accordingly to frame the argument in terms of this trade-off, rather than attributing the effect specifically to k_inj. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      I have a related experimental criticism. The kinetic model presented assumes an exponential decay of free phage, which is a commonly used assumption in the phage literature. Given that the phage types used in this study lyse relatively slowly, it would be good to actually see adsorption curves, in which free phage is measured at different time points between inoculation and lysis. This data would not only provide useful evidence for the kinetic model, but it should also replace what is now in Figure S3, which consists of fitting one experimental point with one line. As it currently stands, Figure S3 is not useful actually misleading.

      We appreciate the reviewer’s point. We agree that adsorption curves, in which free phage is measured at different time points between inoculation and lysis, would provide a stronger basis for evaluating the kinetic model. However, we do not have the resources to perform these additional experiments within the scope of the present study. Following the comments of all reviewers on this point, we have therefore decided to remove Figure S3 to avoid confusion.

      Finally, it is not clear to me why the quantity "Ratio" has been chosen to be presented in Figure 1, rather than the ratio of estimated adsorption rates eta'/eta, which is much more intuitive for a phage study and contains the same information. I would recommend switching to this choice, unless there is a clear rationale for why the quantity "Ratio" is more useful/effective. Showing eta'/eta would also increase the readability of Figure 1, as it would move the y-axis to a logarithmic scale and better visualize values around 1.

      We used “Ratio” in Figure 1 to illustrate the experimental design, controls, and measured quantities directly, as it more transparently reflects the data collected. In the second part of the analysis, where we compare time-independent adsorption rate estimates, we have presented the corresponding values of η′/η as suggested.

      Minor comments:

      (1) Introduction

      Line 31: "... such as nutrient limitation, fluctuating temperatures, and variable energy availability" - if drawing a distinction between energy availability and nutrient limitation, please make explicit what this distinction is. Energy availability seems like a natural consequence of nutrient availability.

      While energy and nutrient availability are often linked in E. coli, they represent distinct physiological constraints. Nutrient limitation refers to the lack of essential biosynthetic precursors such as nitrogen, phosphorus, or amino acids. Energy availability, in contrast, reflects the cell’s ability to generate ATP and reducing equivalents through metabolic processes. For example, under anaerobic conditions, E. coli may have ample nutrients but limited energy production due to the lower efficiency of fermentation compared to aerobic respiration. Thus, energy limitation can occur independently of nutrient limitation.

      (2) Results

      (a) Whole Section: Please label equations.

      All equations have now been labelled in the revised manuscript.

      (b) Lines 105 to 114: As stated in Major Comments, I think the clarity of the paper would be improved by introducing the relative adsorption rate here and dropping the concept of Ratio entirely. However, if the authors wish to use Ratio, I would recommend the following:

      Lines 105 to 109 are confusing to read because of the number of connectives: "... ratio of free viruses from permissive AND resistant hosts respectively TO the free viruses in buffer under energy-depleted AND energy competent conditions". This would be clearer if each quantity were given an algebraic symbol, and RP, RR, and Ratio were defined through formal algebra, rather than mixed mathematical and sentence notation.

      This section has been rewritten for clarity. We now introduce explicit algebraic symbols and define the quantities formally, which removes the ambiguity present in the sentence-only description while retaining the intended meaning.

      The chemical names "arsenate" and "azide" should appear in the body of the text before they appear abbreviated in an equation. Please state at this point that these are both metabolic inhibitors, as it is not immediately clear what role they play or why you are using them.

      The text has been updated to introduce arsenate and azide by name before the abbreviations are used, and we now explicitly note that they act as metabolic inhibitors.

      On line 114, the authors helpfully provide an interpretation of Ratio = 1. It would be useful to provide at the same time interpretations of Ratio >1 and <1, perhaps 2 and 0.5 specifically?

      We have added brief explanations illustrating the interpretation of Ratio values greater than and less than 1, including examples of 2 and 0.5.

      I would consider giving this quantity a more interpretable name than Ratio. This quantity represents how much a bacteriophage preferentially adsorbs to metabolically active cells, so perhaps "Selectivity" or "Adsorption Bias"?

      We intentionally retained the generic term “Ratio”, as this quantity reflects an intermediate experimental measure used to describe the process rather than a newly defined metric. Its purpose is to bridge the experimental observations and the subsequent quantification of effects on the adsorption rate (η).

      (c) Lines 117 to 122: the authors sometimes refer to ratios explicitly, "average ratio of around 1.6" and other times say e.g., "a greater than 3 times increase in viral particles". Using more consistent language (saying "Ratio" every time) would be clearer.

      We have standardized the terminology in this section and now refer to all fold-changes consistently using “Ratio” to avoid ambiguity.

      (d) Figure 1

      Phages λ and T6 look like they have ratios less than 1 for resistant cells? If this is true / if the ratio is statistically significantly below 1, please comment.

      The ratios for λ and T6 are not statistically different from 1. The apparent deviation is within the standard error of the mean. To make this clearer, we have added the corresponding p-values to Table S2 in the Supplementary Information.

      Ratios near 1 are difficult to distinguish from 1, especially in panels A and D. Using a logarithmic scale on the y-axis would make the plots more readable.

      Because the values in these panels are not statistically different from 1, changing to a logarithmic scale would not alter the interpretation. We therefore retained the current axis scaling to reflect that there is no meaningful deviation from 1 in these cases.

      The data corresponding to individual experiments have no error bars. Given that the number of free virions was determined by plaque assay, which carries an intrinsic sampling error, this uncertainty should be reflected in the plots.

      We thank the reviewer for this important comment. Because plaque assays have compound sources of stochastic variation, assigning a per-measurement error bar would risk implying false precision. For this reason, we present the values from each biological replicate directly, and the uncertainty is represented in the statistical summary across replicates. Specifically, for each phage and condition we show the three independent experimental measurements and report the mean along with the standard error of the mean. This approach allows us to represent biological variability without implying a precision that cannot be accurately quantified at the level of single plaque counts.

      Similarly, the average value does show error bars, but it is not stated what these error bars correspond to: standard error in the mean, standard deviation of the sample, or combined uncertainty?

      The caption has been updated to state that the error bars represent the standard error of the mean.

      The resistant bacteria seemed to have ratios close to 1 in all cases. Is this because very few virions adsorbed under both energy conditions?

      Resistance is commonly associated with a lack of a surface receptor for the phage (or generally an entry pathway). We use the resistant bacteria as a control group for the effect of the conditions on adsorption. For resistant bacteria, the Ratio should be 1 since virions do not adsorb under both energy conditions. Any slight variations from 1 should come from sampling errors or small heterogeneity in the population.

      (e) Figure 2

      Please comment on what the error bars here represent. Error bars in Figure 2 A seem to permit negative (or at least zero) values of relative adsorption rate for phages m13 and T6, possibly implying an overestimate of the error? If it is the case that multiple values used to calculate the mean are far apart, possibly showing the values individually through a superimposed swarm plot would be clearer.

      This point is now addressed in the Supplementary Information, where we clarify how the error bars were calculated.

      (3) Discussion

      (a) Line 189: "high metabolic state" is imprecise. Say "energy-competent" to be consistent with earlier language.

      To maintain continuity with earlier terminology, we now include “energy-competent” in parentheses alongside “high metabolic state,” while retaining the original phrasing for readability.

      (b) Figure 3, population level

      Show adsorbed virions physically attached to bacteria, rather than removing them completely from the image, as currently, the implication is that at a high metabolic state, there are fewer virions total, not fewer virions remaining in solution because more are adsorbed. You could go as far as to add a third "after centrifuging" row, showing the adsorbed phages stuck in the pellet and the unadsorbed phages remaining in solution.

      Thank you for this suggestion. Figure 3 has been updated to depict adsorbed virions attached to bacterial cells, clarifying that the decrease represents adsorption rather than loss of total particles. This change improves the accuracy and interpretability of the schematic.

      (4) Methods and Materials

      (a) Figure 5

      The step "estimate cell numbers from OD" appears to follow incubating plates overnight. If the cells you are counting come from the pellet produced by centrifuging 3 steps prior, you could add a fork into the black line connecting the steps, with one branch corresponding to the supernatant and phages, and the other to the pellet and cells?

      Thank you for pointing this out. The order in the figure has been corrected: cell numbers are estimated from OD before overnight incubation. This resolves the confusion without the need for branching in the workflow diagram.

      (a) Line 332

      You allow as much time as possible for adsorption without the possibility of lysis. Did you determine the lysis times / latent periods of these phages through one-step-growth-curves, or use published results, in which case please cite? Having obtained the lysis time by either method, what fraction of the lysis time did you allow for adsorption? Also, please add supplementary tables with lysis times used for the different phages.

      We thank the reviewer for this comment. We used published latent-period values as guides and verified compatibility with our own system when selecting incubation times. We have clarified this in the text and added the relevant citations. We did not use a common fixed fraction of the lysis time for all phages; instead, incubation times were chosen to allow sufficient time for adsorption but not for completion of the first lytic cycle. For λ, productive lytic development was blocked in the host background used, as in Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119. For ϕ80 and T5, we used published latent-period values as guides and verified their compatibility with our own system (De Paepe and Taddei, PLoS Biology 2006, http://dx.doi.org/10.1371/journal.pbio.0040193). M13 is a chronic filamentous phage and therefore does not have a standard lytic latent period; in our host–phage combination, it required more than 1 h before phage release. For T6, we relied primarily on the kinetics observed in our own system, since adsorption was unusually slow for this phage–host pair under our assay conditions. Although literature reports describe shorter T6 latent periods under specific assay conditions (Foster and Johnson, Journal of General Physiology 1951, http://dx.doi.org/10.1085/jgp.34.5.529), this is consistent with published work showing that adsorption and infection kinetics can vary substantially with host background, surface structure, and experimental conditions (Heller and Braun, Journal of Bacteriology 1979, http://dx.doi.org/10.1128/jb.139.1.32-38.1979; Storms et al., Biochemical Engineering Journal 2012, http://dx.doi.org/10.1016/j.bej.2012.02.010).

      (5) Supplementary

      Figure S1

      This data is useful in understanding the main body of the paper, and I think this should form part of a main figure (possibly with the individual experimental data points superimposed over the bars). This could come before or as part of Figure 1?

      We thank the reviewer for this suggestion. We have explored including these data directly in the main figure but found that doing so substantially reduced the readability of the figure, as the underlying table is visually dense. For this reason, we chose to summarize the results in Figure 1 and present the detailed data separately in Figure S1 of the Supplementary Material, along with the Ratio analysis, which more effectively conveys the trends without overloading the main figure.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) L16-18: This sentence could be made more accessible as 'error correction' is not an intuitive term in the phage field.

      We have updated the overall theory section including the terminology. Instead of error correction, we now refer to it as a discrimination process.

      (2) L96-98: Does this potentially indicate a trade-off where evolution for stronger binding cannot evolve at the same time as responsiveness to metabolic activity?

      We agree that this sentence made a stronger evolutionary claim than our data support. Since we only tested four laboratory phages, we cannot conclude that there is an evolutionary trade-off between stronger binding and responsiveness to host metabolic activity. We have therefore removed this sentence to avoid making an unsupported evolutionary interpretation.

      (3) L102: What does 'post-cellular' mean?

      Postcellular supernatant is simply the liquid that remains after cells have been removed. During centrifugation, the cells pellet at the bottom, and the liquid above (which can contain viruses) is the postcellular supernatant.

      (4) L105-107: Worth splitting into two sentences as it is a bit unclear if ratios are built between permissible and resistant hosts or between buffers or both.

      Thank you for the suggestion. We have rewritten this section into two sentences to clarify how the ratios are constructed, and we hope the revised wording improves readability.

      (5) L110-122: Figures S1 and S2 could be referenced here.

      References to Figures S1 and S2 have now been added in this section.

      (6) L137: As P(0) is the viral concentration in buffer, I am assuming that the phage lysate has been diluted in buffer and phages have been added to cultures from the same dilution tube to guarantee equal starting numbers, but I couldn't find this in the methods.

      This clarification has been added to the Methods and Media section of the Supplementary Information.

      (7) L243: It would be worth defining what 'hyperdiffusion' means.

      We have added a brief definition of “hyperdiffusion”.

      (8) L253-256: I do not entirely follow this explanation.

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #3. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (9) L284: Why is k_inj necessarily more sensitive to the metabolic state than k_off? Could membrane changes under stress increase k_off?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: reduced commitment in inactive cells can arise through an increase in k_off, a decrease in k_inj, or both, as long as k_off/k_inj becomes larger in the inactive case. What matters is therefore not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also underlies the trade-off now discussed in the manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become much larger in the inactive case. We have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (10) Figure 1: There seems to be more variation between replicates in phage Lambda than in other phages. Is this caused by receptor number heterogeneity in the population?

      Unfortunately we do not have a way to compare receptor number heterogeneity across the different phage receptors in our experiments. We therefore cannot conclude that the larger variation observed for phage λ is caused by receptor number heterogeneity in the population.

      (11) Figure S1: There seems to be a significant difference between phage Lambda viability in the two buffers - do the authors have an idea where this comes from?

      There is no difference in λ viability between the two buffers. The apparent difference in the figure is due to sampling variability.

      (12) Figure S3: Last sentence of the legend probably shouldn't say 'upper'.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      Reviewer #3 (Recommendations for the authors):

      (1) The text reads as incomplete in some places. Can the authors please provide clarifications on the following points?

      (1.1) Lines 235-256: How do the authors draw a conclusion that "a phage can detect host metabolic status" from a study that used purified LamB receptors (i.e., no live cells with any metabolism) extracted from two different bacterial species (i.e., not a difference in metabolic states)?

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #2. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (1.2) Line 270, in the abstract, and in the caption of Figure 4: The authors described the model using terms such as "an error-correction mechanism" or "standard error correction", but there is little explanation. Can the authors clarify what kind of "error" is discussed here, and how it is "corrected"? In the "standard error correction" model, what determines which method of correction is "standard"? If "error correction" is a standard term in phage-bacterial interaction modeling, please provide references.

      We agree with the reviewer that our use of the term error correction was not appropriate in this context. The proper term is discrimination process rather than error correction. We have now corrected this terminology throughout the manuscript and clarified the underlying logic in the relevant sections.

      (1.3) Line 301: The authors speculated that phage T5 is "better suited to ecological niches", but I am not sure how that is consistent with their data showing T5 is more rampant, that they infect both energy-competent and energy-depleted cells, not just depleted cells. Why "niches", and why are T5 better suited to environments "where energy-limited cells dominate", not just any environment?

      We agree that this point was not stated clearly enough. What we intended to convey is that T5 would be at a net disadvantage in a niche containing a mixture of energy-competent and energy-deficient hosts. We have updated the main text accordingly.

      (1.4) Line 303, and related to point 6.3. above: Phage λ can also infect and replicate in "starved bacterial cells" (shown in Kourilsky 1974 and Geng et al. 2024, both of which were cited in this manuscript). How do the authors reconcile these reports with the discussion point in line 303, and their data that only phage T5, but not λ, shows insensitivity to the host metabolic state?

      Our data do not imply that phage λ is unable to infect starved bacteria. As shown in Kourilsky (1974) and Geng et al. (2024), λ can indeed infect and replicate in nutrient-limited cells. Our results specifically indicate that λ infection under starvation proceeds with a reduced adsorption rate, while T5 maintains the same adsorption rate even when the host is starved. Thus, our conclusion is that T5 is insensitive to the host metabolic state at the level of adsorption, whereas λ is not. We acknowledge that the wording in line 303 may have unintentionally led to confusion, and we have revised this part of the text to avoid that.

      (2) The following comments relate to the text and figures in the manuscript. There are many places in the manuscript that could use fine proofreading and copy-editing for clarity and consistency. For example:

      (2.1) If I understand it correctly, the equation in between lines 109 and 110 should be clarified using terms such as "Free viral particles after mixing with bacteria in Arsenate and Azide" and "Free viral particles in bacteria-free buffer with Arsenate and Azide". As it stands, it is not clear which terms correspond to conditions where bacteria are present.

      The equation has been updated to explicitly indicate which terms refer to mixtures containing bacteria and which refer to bacteria-free controls, so that the correspondence between conditions is now clear.

      (2.2) Equations in between line 276 and 283, and elsewhere: Some concentration terms are enclosed in brackets ("[BP]"), while most are not.

      This notation has been clarified. We now use “[PB]” specifically to denote the transient phage–bacterium complex, distinguishing it from the product P⋅B. All other concentration terms are written without brackets for consistency.

      (2.3) Figure 4 and in equations: "BP" or "PB"?

      The notation has been made consistent throughout; we now use “PB” exclusively to denote the phage–bacterium complex.

      (2.4) Line 284 and line 286: The "inj" in "k_inj" is sometimes italicized, sometimes not.

      The notation has been standardized so that k_inj is now formatted consistently throughout the manuscript, without italicizing “inj.” Also we have replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (2.5) Figure 5: Was the step "Estimate cell numbers from OD" really performed on the next day after the experiment (i.e., >12 hours after infection and phage plating), not immediately after cell washing?

      Thank you for pointing this out. The figure has been updated to reflect the correct order of steps: cell numbers are estimated from OD immediately after washing, followed by overnight incubation of the plates.

      (2.6) Figure S1: As it stands now, the x-axis of each panel can be read either as "Permissive, Resistant bacteria, Buffer" (missing "bacteria" for the first pair of bars), or "Permissive (bacteria), Resistant (bacteria), Buffer (bacteria)" (extra "bacteria" for the last pair of bars).

      The intended interpretation is the second one (permissive bacteria, resistant bacteria, buffer).

      (2.7) Figure S3: The panel letters "A" and "B" are missing in the figure. Also, it is not clear why the legend for the five phages and the legend for the measurement times are not combined.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      (2.8) Strain table in the Methods and Materials: Please write genotypes with italicization, and consistently indicate mutations and deletions with the minus sign superscript or the Δ prefix. Also, for the S3222 strain: Is it really the entire Mal regulon mutated ("Mal-"), or just lamB-? In Brown et al. 2022, it was only the latter.

      Genotypes have been reformatted with consistent notation. For S3222, the correct designation is Mal-, as in the SI of Brown et al. 2022. In this case, Mal- is intended as a phenotypic designation rather than a specific genotype, and we have therefore formatted it accordingly, i.e. neither italicized nor written in lower case.

    1. eLife Assessment

      This study makes a valuable contribution by broadening the range of eukaryotic model systems and establishing Blastocystis, the most prevalent microeukaryote in the human gut, tractable for reverse-genetics investigations. The presented imaging data are convincing and informative, although confirmation by molecular methods would further strengthen the study. The work should interest readers studying host-microbe interactions in the human gut, as well as those developing new systems for eukaryotic research.

    2. Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in Blastocystsis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did - so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      Comments on revised version.

      The authors have revised their manuscript to clarify that molecular analyses have not yet occurred and have resolved the technical/publication issues with the figures. I look forward to seeing these tools used in future publications to answer important questions in Blastocystsis research.

    3. Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide a solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to visual presentation of the data would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested "23 promoter-terminator pairs" to better reflect that only promoters varied.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, the part A of Figure 2 is split into two panels with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of the part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible and it draws attention to a distracting design feature rather to the data themselves.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      (4.2) In the parts B-D, the legend should explain more clearly what each image shows and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images as the current distortions complicate interpretation.

      Comments on revised version.

      The revised version provides sufficient clarity and appropriate visual presentation. Some confusion evidently arose due to my misunderstanding, so I thank the authors for their comprehensive clarifications and patience.

    4. Author response:

      The following is the authors’ response to the original reviews.

      In revising the manuscript, we have focused on three main priorities raised during review: (1) improving precision around evidential claims, particularly concerning vector maintenance and P2A-mediated protein separation; (2) substantially improving figure quality, accessibility, and legend clarity; and (3) correcting inconsistencies and expanding methodological detail where requested.

      This study was intended as a foundational genetic toolkit and methodological framework for Blastocystis ST7-B, establishing practical workflows for DNA delivery, endogenous regulatory-element benchmarking, antibiotic-selected recovery, clonal propagation, and reporter-based analysis in a genetically challenging anaerobic microbial eukaryote. The central evidence presented is therefore functional in nature: reproducible transgene delivery, selectable recovery and propagation of colony-derived transgenic lines, and detectable reporter expression using multiple anaerobic-compatible reporter systems.

      We agree with the reviewers that several additional experiments, including Western blot analysis of P2A-containing constructs, outward-facing PCR, plasmid rescue assays, and selection-withdrawal experiments, would further strengthen the mechanistic interpretation of the system and help distinguish episomal persistence from genomic integration. We have therefore revised the manuscript throughout to clearly separate what is directly demonstrated from what remains a plausible working interpretation or important future direction.

      Importantly, the revised manuscript no longer presents episomal maintenance or complete P2A-mediated protein separation as demonstrated conclusions. Instead, these are now discussed explicitly as unresolved mechanistic questions requiring future molecular analysis. Nevertheless, the central methodological conclusion remains unchanged: stable selectable transgene expression, recovery of colony-derived transgenic lines, and reporter-positive Blastocystis ST7-B transformants can now be reproducibly obtained.

      Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in the Blastocystis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did, so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      We thank the reviewer for the positive assessment of the manuscript and for recognising both the technical difficulty and broader utility of establishing genetic tools in Blastocystis and other experimentally challenging microbial eukaryotes. We also appreciate the reviewer’s identification of the main evidential limitations in the original manuscript, particularly regarding vector maintenance and P2A-mediated protein separation.

      First, regarding the question of vector topology and the interpretation of episomal maintenance.

      We agree that the original manuscript presented episomal persistence too strongly relative to the evidence currently available. We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working model rather than a directly demonstrated conclusion.

      The transfection system used here was adapted from Li et al. (2019), including use of the pXS2-P<sub>Legumain</sub>-derived plasmid framework. Importantly, the construct used in the present study does not contain the original Trypanosoma brucei tubulin-targeting region associated with homologous integration in the original pXS2 system. Complete plasmid sequencing confirmed that the constructs function here as heterologous expression plasmids carrying Blastocystis ST7-B regulatory elements and transgenes. While this does not demonstrate episomal persistence, it also means that genomic integration cannot be inferred from the historical pXS2 vector architecture alone.

      We further note that comparative genomic analyses by Gentekaki et al. (2017) suggest that Blastocystis lacks components of the canonical non-homologous end-joining (NHEJ) machinery, implying that homologous recombination is likely to represent the principal route for double-stranded DNA repair. Because the constructs used here did not contain Blastocystis homology arms, there is currently no obvious mechanism favouring targeted homologous integration. Nevertheless, we fully agree that genomic integration cannot presently be excluded.

      To reflect this appropriately, the revised manuscript now explicitly separates the demonstrated functional outcomes from unresolved mechanistic questions concerning vector maintenance. We also identify several future approaches that would help distinguish episomal persistence from genomic integration, including outward-facing PCR, plasmid rescue followed by full plasmid sequencing, Southern blotting, FISH, selection-withdrawal experiments, and long-read sequencing approaches.

      We have revised the manuscript throughout to remove statements implying demonstrated episomal maintenance and now present episomal persistence only as a plausible working interpretation.

      In the Methods section under Cloning, the following text has been added:

      Lines 202–206: “The constructs used in this study were derived from the pXS2-P<sub>Legumain</sub> vector described by Li et al. (2019), which adapted a heterologous expression-vector backbone for transient plasmid-based expression in Blastocystis ST7-B. Here, the same molecular backbone was used as a plasmid scaffold carrying Blastocystis-derived regulatory elements and transgenes.”

      In the Discussion, the following text has been added/edited:

      Lines 665–673: “The molecular maintenance state of the introduced constructs remains unresolved: episomal maintenance is a plausible working model, but genomic integration cannot be formally excluded. The constructs used here lack Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components (Gentekaki et al., 2017), making targeted integration by standard repair routes unlikely but not impossible. Direct assays such as outward-facing PCR, plasmid rescue followed by full plasmid sequencing, FISH, or selection-withdrawal experiments will be required to distinguish episomal persistence from integration.”

      Second, regarding P2A-mediated protein separation.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. We have therefore revised the manuscript to avoid implying that complete protein-level separation was directly demonstrated.

      The revised manuscript now states only what is directly supported by the current data: that P2A-containing bicistronic constructs supported antibiotic-selected recovery of transgenic lines together with detectable downstream reporter expression. The microscopy data therefore support functional downstream reporter expression, but do not by themselves exclude residual uncleaved fusion products.

      We selected P2A because it is a compact and well-characterised peptide with high reported separation efficiency across multiple eukaryotic systems, including microbial eukaryotes. However, we agree that P2A performance can be context-dependent, and we now explicitly identify biochemical validation of P2A cleavage efficiency as an important future direction.

      Importantly, these revisions do not alter the central methodological conclusion of the study, namely that selectable transgene expression, propagation of reporter-positive lines, and recovery of colony-derived Blastocystis ST7-B transformants can now be reproducibly achieved.

      Text inserted in the Results:

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.”

      Lines 403–404: “However, protein-level separation was not directly tested, and the extent of any residual uncleaved fusion product remains unresolved.”

      Text inserted in the Discussion:

      Lines 619–629: “P2A was selected because it is a well-characterised peptide with high reported separation efficiency in human cell lines, zebrafish embryos, and mice (Kim et al., 2011). It also has precedent across microbial eukaryotes, including the protest Dictyostelium discoideum (Zhu et al., 2023), the fungi Aspergillus niger (Schuetze and Meyer, 2017) and Ustilago maydis (Müntjes et al., 2020), and the apicomplexan parasites Toxoplasma gondii (Markus et al., 2019) and Plasmodium falciparum (Dans et al., 2024). However, P2A performance is context-dependent, and the evidence presented here is functional rather than biochemical. P2A-containing constructs support antibiotic-selected recovery and downstream reporter expression in Blastocystis ST7-B, but ribosomal skipping efficiency and any residual uncleaved product will require direct protein-level validation.”

      Reviewer #1 (Recommendations for the authors):

      (1) Please could you show a Western blot to confirm if P2A has worked? It could be that the proteins are being expressed as a polyprotein.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. This is an important point, and we have revised the manuscript accordingly to avoid implying that complete protein-level separation was directly demonstrated.

      The current study was designed as a first-generation functional genetic toolkit for Blastocystis ST7-B, focused primarily on establishing reproducible workflows for selectable transgene expression, reporter recovery, and propagation of transgenic lines in this experimentally challenging anaerobic microbial eukaryote. The toolkit is therefore validated here through functional outcomes, including antibiotic-selected survival, stable propagation through extended passaging (>15 passages) and cryopreservation, and detectable reporter fluorescence above wild-type autofluorescence.

      P2A was selected because it is a compact and well-characterised peptide with high reported ribosomal skipping efficiency across multiple eukaryotic systems, including microbial eukaryotes, as discussed above. Nevertheless, we fully agree that direct biochemical validation would strengthen the mechanistic interpretation of the bicistronic system in Blastocystis ST7-B. We therefore now explicitly identify Western blot analysis, ideally using epitope-tagged upstream and downstream products, as an important future direction for quantitative assessment of P2A cleavage efficiency and any residual uncleaved fusion products.

      Relevant manuscript revisions are described above under the general response to Reviewer 1.

      (2) Something has gone wrong with figure formatting. Figure 2 is nearly illegible and I cannot read the text in section A. Sections B, C, and D have lost their labels and are fuzzy and surrounded by black. A similar issue affects Figure 3. Everything is just black with a few cells. It is illegible when printed.

      We thank the reviewer for highlighting these presentation issues and agree that the submitted figure quality significantly impaired readability and interpretation. The problems appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, contrast, and panel labelling in the review PDF.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved typography, panel separation, colour scaling, and legend clarity throughout. In response to additional reviewer suggestions, individual data points have now been added to Figures 2B and 2C to improve transparency and interpretability of the underlying data distributions.

      Figures 2 and 3 have been replaced with fully revised high-resolution versions with improved panel labelling, accessibility, typography, and figure legends.

      (3) The data from Figure 2B would be better placed in Table 1 with a column for robust/moderate/intermediate/weak/very weak. This would be much easier for the reader.

      We thank the reviewer for this helpful suggestion. We believe the comment refers to the promoter activity data shown in Figure 2A rather than the voltage optimisation data in Figure 2B. To improve readability and accessibility of these data, we have revised Figure 2A extensively to make the promoter activity tiers more legible and easier to interpret directly from the heat map and accompanying box plots.

      We considered incorporating simplified activity classifications into Table 1. However, activity patterns were construct-specific rather than simply locus-specific. In several cases, multiple promoter fragments derived from the same locus produced substantially different reporter outputs, and activity did not scale monotonically with promoter fragment length. We therefore felt that assigning a single categorical activity label at the locus level would oversimplify the dataset and reduce the construct-level resolution that is central to the toolkit value of the study.

      Instead, we addressed the reviewer’s concern by substantially improving the presentation and readability of Figure 2A, allowing readers to identify robust, moderate, intermediate, weak, and very weak expression constructs more directly while preserving the underlying construct-specific information.

      Figure 2A has been revised to improve clarity, accessibility, and legibility of the promoter activity tiers, allowing construct-level expression classes to be interpreted more directly from the heat map and accompanying boxplots.

      (4) How do you know if the constructs are maintained as episomes? Have you done an outward-facing PCR?

      We agree that direct molecular evidence distinguishing episomal persistence from genomic integration is currently lacking, and we appreciate the reviewer highlighting this important limitation. We have therefore revised the manuscript throughout to avoid presenting episomal maintenance as a demonstrated conclusion and now describe it only as a plausible working interpretation based on the current evidence and vector design.

      We have not performed outward-facing PCR in the present study. As discussed in the general response above, we now explicitly identify outward-facing PCR, plasmid rescue followed by full plasmid sequencing, selection-withdrawal assays, FISH, and long-read sequencing approaches as important future directions for resolving the molecular maintenance state of the constructs.

      The revised manuscript now clearly separates the demonstrated functional outcomes, including selectable transgene expression, recovery of colony-derived transgenic lines, and stable reporter-positive propagation under selection, from the unresolved mechanistic question of vector topology.

      This issue has been addressed throughout the revised manuscript, including in the Methods and Discussion sections, where episomal maintenance is now presented as a plausible but unconfirmed interpretation rather than a demonstrated conclusion.

      Minor Comments

      Line 66: is this one to two billion individuals with Blastocystis, or one to two billion Blastocystis cells per gut?

      The intended meaning was colonised individuals globally. We agree that the original phrasing was ambiguous and have corrected it for clarity.

      Lines 66–67 revised to: “…microorganisms in the human gut, and is estimated to colonise approximately one to two billion people globally (Scanlan and Stensvold, 2013).”

      Line 148: Supplier of IMDM?

      The supplier information was already present in the original manuscript as IMDM L0191 (Biowest).

      No additional manuscript change required.

      Line 157: Who annotated the dataset, the 2017 paper or the present study?

      The dataset annotation derives from Armengaud et al. (2017). We agree that the original wording was unclear and have revised this section substantially to improve clarity regarding the rationale and workflow used for promoter and terminator candidate selection.

      “The relevant Methods section has been extensively revised for clarity and expanded detail” (Lines 156–189).

      Line 166: Who predicted the 3′ UTR, the 2017 paper?

      This information derives from the NCBI annotation associated with the Blastocystis ST7-B genome based on Denoeud et al. (2011). This has now been clarified in the Methods section.

      Clarified in revised Methods section.

      Line 237: How long did it take in days?

      Approximately 2 days.

      Line 270 revised to: “…turned yellow without drug treatment, usually within 2 days post-transfection.”

      Line 325: Typo, missing gap between Figure and 1A.

      Corrected in revised manuscript.

      Reviewer #2 (Public review):

      This manuscript presents a substantial technical advance for the genetic manipulation of Blastocystis by establishing an integrated workflow for stable episomal transgenesis, antibiotic selection, clonal recovery, and reporter-based imaging in the ST7-B subtype. The study is particularly valuable because it combines multiple previously fragmented approaches into a coherent and practically applicable toolkit, including endogenous regulatory elements, optimized electroporation conditions, selectable markers, and anaerobic compatible fluorescent reporters. This methodological work greatly expands the molecular toolbox and future studies focused on both basic and infection biology can now build on the ability to express and localize proteins in fixed as well as live cells.

      The microscopy data are convincing and clearly demonstrate functional reporter expression and successful recovery of stable transgenic lines. Nevertheless, because this is primarily a methodological paper, the study would be further strengthened by the inclusion of Western blot validation of reporter expression and bicistronic constructs. In particular, biochemical analysis of the P2A-containing constructs would help assess the efficiency of ribosomal skipping and exclude the possible presence of uncleaved fusion proteins, thereby providing stronger support for the interpretation of the imaging data and the functionality of the expression system.

      We thank the reviewer for this thoughtful and positive assessment of the manuscript and for recognising the value of integrating previously fragmented approaches into a coherent and practically usable genetic toolkit for Blastocystis ST7-B. We particularly appreciate the reviewer’s recognition that the system expands the currently available molecular toolbox for both cell biological and infection-related studies in this experimentally challenging anaerobic microbial eukaryote.

      We also appreciate the reviewer’s comments regarding biochemical validation of the P2A-containing bicistronic constructs. We agree that Western blot analysis would strengthen the mechanistic interpretation of the reporter system by directly assessing ribosomal skipping efficiency and the possible presence of residual uncleaved fusion products. In response, we have revised the manuscript throughout to ensure that the conclusions remain appropriately evidence-based and do not imply that complete protein-level separation was directly demonstrated.

      The revised manuscript now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, stable propagation of reporter-positive lines, and detectable downstream reporter expression, from unresolved mechanistic questions concerning P2A cleavage efficiency and vector maintenance state. We now also identify biochemical validation of P2A-mediated protein separation as an important future direction for further refinement of the system.

      Relevant manuscript revisions addressing these points are described above under the response to Reviewer 1.

      Reviewer #2 (Recommendations for the authors):

      The quality of images could be better. The figures lacked resolution — possibly a conversion artefact.

      We agree that the figure quality in the submitted review PDF significantly reduced readability and visual interpretation. The issues appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, typography, panel labelling, and contrast rendering.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved panel separation, typography, colour scaling, contrast settings, and figure legends to improve accessibility and interpretability both on screen and in print. In addition, the export workflow and file formatting have been updated to improve compatibility with journal production requirements and reduce the likelihood of compression-related rendering artefacts during manuscript compilation.

      Figures 2 and 3 have been replaced with revised high-resolution versions with improved typography, panel labelling, contrast settings, and accessibility.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for the propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design, to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to the visual presentation of the data, would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      We thank the reviewer for this important point and agree that the original manuscript presented episomal persistence too strongly relative to the currently available evidence. In particular, we agree that the title and several sections of the manuscript implied a level of mechanistic certainty that was not directly demonstrated.

      We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working interpretation rather than a demonstrated conclusion. The revised text now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, recovery and propagation of colony-derived transgenic lines, and stable reporter-positive maintenance under selection, from the unresolved mechanistic question of vector topology.

      As discussed in our response to Reviewer 1, the constructs used here do not contain Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components, making targeted integration by standard repair routes less strongly supported mechanistically, although genomic integration cannot presently be excluded.

      We agree that selection-withdrawal experiments would provide useful additional evidence regarding construct persistence and have now explicitly identified such assays, together with outward-facing PCR, plasmid rescue, FISH, and long-read sequencing approaches, as important future directions for resolving the molecular maintenance state of the transgenes.

      The manuscript has been revised throughout to remove wording implying demonstrated episomal maintenance. Episomal persistence is now discussed only as a plausible working interpretation pending direct molecular validation.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      We thank the reviewer for this careful reading and for identifying this inconsistency. We agree that the distinction between candidate loci and successfully generated promoter constructs was not sufficiently clear in the original manuscript and could create uncertainty regarding the dataset.

      The original candidate set comprised 14 loci selected for promoter and terminator discovery. However, only 11 loci yielded successfully cloned and experimentally tested promoter constructs. The remaining three loci were retained in Table 1 for completeness and transparency, as repeated cloning attempts were unsuccessful despite two independent efforts.

      We have revised the manuscript to make this distinction explicit and to clarify that the reported NanoLuc benchmarking experiments were ultimately performed using constructs derived from 11 successfully cloned loci.

      Lines 361–364: “To expand the available regulatory parts, we screened 23 NanoLuc reporter constructs containing putative endogenous promoter–terminator pairs from 11 of 14 candidate loci; three loci could not be cloned after two independent attempts and are indicated in Table 1.”

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested “23 promoter-terminator pairs” to better reflect that only promoters varied.

      We thank the reviewer for the opportunity to clarify this point. We agree that the original presentation may have created the impression that promoter and terminator regions were independently varied and benchmarked, whereas the experimental design was primarily focused on construct-level comparison of endogenous regulatory modules.

      As described in the Methods, each construct contained a candidate endogenous upstream promoter region together with the corresponding endogenous downstream terminator region derived from the same locus. For consistency and to keep the cloning and screening strategy experimentally tractable, a fixed 500 bp downstream terminator fragment was used for each locus rather than systematically varying terminator length or independently testing terminator activity.

      We therefore retain the description “endogenous promoter–terminator pairs,” since each construct contains both endogenous upstream and downstream regulatory regions from the same genomic locus. However, we agree that the assay was not designed to independently dissect promoter versus terminator contributions to reporter output. We have revised the manuscript accordingly to make this distinction explicit and avoid ambiguity regarding the scope of the benchmarking analysis.

      Lines 365–368: “Each construct paired a candidate upstream promoter region with the corresponding downstream terminator region from the same locus, defined here as the native 500 bp sequence immediately downstream of the stop codon. Where multiple promoter lengths were tested for the same locus, the terminator fragment was kept constant (Table 1; Figure 1A).”

      This design allowed construct-level benchmarking of paired promoter–terminator modules but did not test promoter strength or terminator activity independently.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      We thank the reviewer for these important methodological points. We agree that the original manuscript did not sufficiently clarify the transient nature of the NanoLuc benchmarking assay or the rationale underlying the assay design and timing.

      The promoter benchmarking assay was designed as an early transient-expression screen adapted from the NanoLuc-based workflow of Li et al. (2019), with modifications, rather than as a stable-maintenance assay. No selectable marker was included because the objective was to compare relative early reporter output across constructs shortly after DNA delivery, before prolonged culture effects became dominant.

      The 16–18 h post-electroporation time point was selected based on the NanoLuc expression kinetics reported by Li et al. (2019) and empirical optimisation during assay development. This window allowed robust transient reporter detection while limiting confounding effects arising from prolonged plasmid loss, differential outgrowth, variable recovery, or later culture-level changes.

      We agree that, in the absence of selection, the observed NanoLuc signal cannot be interpreted as an absolute measure of promoter strength independent of DNA uptake efficiency, early plasmid retention, or post-transfection recovery dynamics. We have therefore revised the manuscript to clarify that Figure 2A reports relative transient reporter output under standardized early post-transfection conditions rather than isolated promoter activity alone.

      We now also explicitly state that the data shown in Figure 2A derive from three independent electroporation experiments per construct, each assayed in technical duplicate.

      Lines 241–248: “Promoter–terminator activity was assessed 16–18 h after electroporation using a transient NanoLuc assay adapted from Li et al. (2019), with modifications. This early time point was selected to capture reporter output within the transient-expression window after DNA delivery, before prolonged plasmid loss, differential outgrowth, or culture-level changes could dominate the readout. Because the constructs did not contain a selectable marker, the measured NanoLuc signal reflects early transient reporter output rather than promoter strength independent of DNA uptake, early plasmid retention, or post-transfection recovery.”

      Additional clarification added to Figure 2 legend stating that measurements derive from three independent electroporation experiments, each assayed in technical duplicate.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, part A of Figure 2 is split into two panels, with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible, and it draws attention to a distracting design feature rather than the data themselves.

      We thank the reviewer for these detailed comments regarding figure design and visual interpretation. We agree that the original presentation of Figure 2 introduced unnecessary visual ambiguity through inconsistent colour usage, the use of a diverging colour scale for strictly positive values, and poor readability associated with the dark background and low-resolution export.

      In response, Figure 2 has been extensively redesigned to improve clarity, accessibility, and interpretability. The previous blue–white–red diverging heatmap has been replaced with a sequential colour palette appropriate for positive-only expression data. Boxplot colouring has also been simplified and harmonised with the heatmap scheme to avoid implying unsupported categorical distinctions. In addition, panel organisation, typography, scaling, and legend structure have all been revised to improve readability and reduce ambiguity regarding interpretation of the plotted values.

      We also agree that the black background distracted from the data presentation and impaired visibility of panel labels and image boundaries. The revised figures therefore use white backgrounds together with clearer panel separation and improved label visibility throughout.

      Figure 2 has been completely reformatted using a sequential colour scale in panel A, simplified and harmonised boxplot colouring, larger typography, improved panel separation, revised legends, and white backgrounds throughout. Corrected high-resolution source figures have been provided.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      We thank the reviewer for this helpful comment regarding figure composition and panel separation. We agree that the original presentation made it difficult to distinguish intentional panel boundaries from imaging artefacts, particularly in the low-resolution review PDF generated during manuscript compilation.

      To improve clarity, Figure 3 has been reformatted using white backgrounds, clearer panel spacing, and more explicit separation between individual snapshots and imaging channels. High-resolution source images have also been provided to ensure that fluorescence patterns, image boundaries, and panel organisation remain clearly interpretable both on screen and in print.

      Figure 3 has been reformatted with improved panel separation, white backgrounds, clearer image boundaries, and revised high-resolution source figures.

      (4.2) In parts B-D, the legend should explain more clearly what each image shows, and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      We thank the reviewer for these helpful suggestions regarding figure annotation and legend clarity. We agree that the original presentation did not sufficiently explain the composition of the imaging panels, particularly under the low-resolution conditions of the review PDF.

      To improve interpretability, the revised Figure 3 now includes clearer panel organisation, improved annotations, and expanded figure legends explicitly identifying the individual imaging channels and staining conditions shown in each subpanel. The leftmost panels in parts B–D are now more clearly identified in both the figure and legend, together with the corresponding fluorescence or staining conditions used in each experiment.

      As mentioned in the Methods sections we used Hoechst 33342 to visualise DNA; but we agree that the Hoechst 33342-labelled structures required additional clarification. The revised legend section now explains that multiple Hoechst 33342-positive regions are commonly observed because Blastocystis cells can contain multiple nuclei depending on cell stage and subtype-specific morphology.

      In addition, high-resolution source images have been provided to ensure that fluorescent signals, panel boundaries, and imaging features remain clearly interpretable both on screen and in print.

      Figure 3 legends and annotations have been revised to clarify imaging channels, staining conditions, and panel organisation. The figure caption was also edited to include: “DNA was visualised using Hoechst 33342. Most cells contained two nuclei, and smaller Hoechst 33342-positive signals consistent with mitochondrial DNA were also observed in some instances.”

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images, as the current distortions complicate interpretation.

      We thank the reviewer for this important observation and agree that the apparent morphological differences between the UnaG/smURFP and SNAP-tag panels required additional clarification.

      The images shown for the different reporter systems were acquired under different imaging conditions and microscope configurations following an instrument-related issue during part of the imaging workflow, as noted in the Methods section. As a result, direct visual comparison of cell morphology between reporter systems is not appropriate. The primary purpose of these panels is instead to demonstrate reporter detectability, live-cell labelling capability, and the characteristic fluorescence patterns obtained with the different anaerobiosis-compatible reporter systems.

      In particular, the SNAP-tag panels were included to demonstrate successful live-cell labelling without permeabilisation together with the expected increase in fluorescence signal at higher substrate concentrations, rather than to support quantitative comparison of cell morphology across imaging conditions.

      We considered replacing the affected images. However, equivalent replacement datasets acquired under directly comparable conditions are not currently available. We have therefore retained the original images but revised the figure legend to clarify the intended interpretation and limitations of these panels explicitly.

      Figure 3 legend revised to include:

      “Because images for the different reporter systems were acquired under different imaging conditions, they are presented to demonstrate reporter detectability and labelling pattern and should not be used for quantitative comparison of cell morphology across reporter systems.”

      Reviewer #3 (Recommendations for the authors):

      The reader may find the current order confusing starting with construct design before testing which drug to use for selection. The narrative would work better if it started with antibiotic selection as the first logical step for generating stable cell lines.

      We thank the reviewer for this thoughtful suggestion regarding narrative structure and agree that multiple organisational strategies are possible for presenting a methodological workflow of this type.

      We considered reorganising the Results section to begin with antibiotic selection and drug sensitivity profiling. However, we ultimately retained the overall structure because the manuscript is organised as a toolkit-development framework rather than as a strictly chronological experimental protocol. The Results therefore begin with regulatory-element discovery and construct design, which form the conceptual and experimental foundation of the toolkit, before progressing to DNA delivery optimisation, drug sensitivity profiling, clonal recovery, and reporter validation.

      We felt that this structure most clearly reflects the dependency relationships within the system: regulatory elements are required before constructs can be assembled, constructs are required before electroporation conditions can be evaluated, and selectable constructs are required before stable selection and clonal recovery can be meaningfully assessed.

      (2) The text states that the screen 'focused on the 1,000 most abundant proteins to establish a preliminary library capable of supporting varying levels of transcription.' Since the genome has ~6,000 protein-coding genes, the top 1,000 cover the most abundant proteins — not a wide expression range.

      We thank the reviewer for this important clarification. We agree that the original wording could incorrectly imply that the screen was intended to sample broadly across the full transcriptional range of the Blastocystis genome. This was not the case, and we have revised the manuscript accordingly.

      Our strategy was instead designed to enrich for candidate loci with a higher prior likelihood of supporting detectable transgene expression. Because no genome-wide promoter map, transcription start site dataset, or experimentally validated regulatory annotation was available for Blastocystis ST7-B at the inception of this work, we used the abundance-ranked Blastocystis ST4-WR1 proteomic dataset of Armengaud et al. (2017) as a practical starting point for candidate discovery.

      Importantly, the Blastocystis ST4-WR1 proteome is highly skewed, with 193 proteins contributing approximately 50% of the detected proteome and the 13 most abundant proteins contributing approximately 10% (Armengaud et al., 2017). We therefore selected the top 1,000 proteins not as a representation of the genome-wide expression range, but as a proteomics-guided enrichment strategy to identify loci more likely to contain active endogenous regulatory regions suitable for initial toolkit development.

      We have revised the relevant Methods section substantially to clarify both the rationale and the workflow used for candidate selection, homolog identification, and promoter/terminator definition.

      The Methods section (Lines 156–189) has been extensively revised to clarify the rationale underlying candidate regulatory-element selection. The revised text now explicitly states that the strategy was designed to enrich for likely active loci for toolkit development rather than to systematically survey the full range of promoter strengths across the Blastocystis genome.

      Additional methodological detail has also been added regarding:

      Use of the Armengaud et al. (2017) proteomic and proteogenomic datasets,

      Homolog identification in Blastocystis ST7-B,

      Locus selection criteria,

      Promoter boundary definition,

      And operational definition of candidate terminator regions.

      (3) The Methods contain an inconsistency: cells were left in 0.5 mL, then 1 mL was added, but then only 0.5 mL is apparently used for transfection. What happened to the 1 mL?

      We thank the reviewer for identifying this ambiguity in the transfection workflow description. The apparent inconsistency arose because the protocol description moved from bulk cell resuspension to preparation of individual electroporation reactions without explicitly stating how the intermediate suspension was used.

      After washing, approximately 0.5 mL of cytomix buffer remained above the pellet, and 1 mL of complete cytomix buffer was then added to generate an approximately 1.5 mL cell suspension. Cells were counted from this pooled suspension, after which the volume corresponding to 5 × 10<sup>7</sup> cells was transferred into each individual electroporation reaction. Following addition of DNA, each electroporation reaction was adjusted to a final volume of 500 µL with complete cytomix buffer. The remaining cell suspension was retained for additional transfections or control reactions.

      We agree that the original wording could be misinterpreted and have revised the Methods section to clarify the sequential handling steps more explicitly.

      Lines 225-229 revised to read: “The resulting approximately 1.5 mL pooled cell suspension was used for total viable cell counting using a hemacytometer.”

      “After counting, the volume corresponding to 5 x 10<sup>7</sup> cells was transferred to each electroporation reaction and combined with 25 µg of plasmid DNA. The total electroporation volume was adjusted to 500 µL with complete cytomix buffer.”

      (4) Figures 2 and 3 are too low-resolution for the font size used and for clearly viewing the microscopy images.

      We thank the reviewer for highlighting these readability issues. As noted in our responses above regarding Figures 2 and 3, the low-resolution appearance primarily resulted from manuscript compilation and PDF export artefacts affecting typography, image rendering, and panel clarity in the review version.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions featuring improved typography, panel labelling, contrast, accessibility, and image clarity for both on-screen viewing and print reproduction.

      Revised high-resolution versions of Figures 2 and 3 have been provided as described above. No additional manuscript changes were required beyond the figure revisions already outlined.

      (5) Figure 4 is confusing because the left and right panels appear inconsistent, with much higher concentrations required for growth inhibition in the culture-based assay than the resazurin assay indicated. The rationale for the resazurin assay should be explained, and the complete growth inhibition (CGI) concentration should be highlighted in the right panel.

      We thank the reviewer for highlighting this potential source of confusion. We agree that the distinction between the two assay endpoints was not sufficiently emphasised in the original figure presentation and legend.

      The apparent discrepancy arises because the two assays measure different biological endpoints under different assay conditions. The resazurin assay was used to estimate IC<sub>50</sub> values, corresponding to the concentration at which metabolic activity was reduced by approximately 50% under the assay conditions. In contrast, the small-culture assay was designed to determine complete growth inhibition (CGI), defined operationally as the concentration at which no detectable culture outgrowth occurred after incubation, using phenol red acidification as a culture-level readout.

      Because these assays measure partial metabolic inhibition versus complete suppression of detectable culture outgrowth, the corresponding concentration ranges are not expected to coincide directly. The higher concentrations observed in the right-hand panels therefore reflect the more stringent endpoint associated with complete growth inhibition rather than inconsistency between the assays.

      We agree that this distinction should have been explained more clearly in the original manuscript. We have therefore substantially revised the Figure 4 legend to clarify the rationale underlying both assays, explicitly distinguish IC<sub>50</sub> and CGI endpoints, and explain how the CGI values were used to guide subsequent antibiotic selection conditions for Blastocystis ST7-B transformants. The CGI transition range has also been made more visually explicit in the revised figure presentation.

      Figure 4 caption revised to: “Antibiotic potency and selection-window determination in Blastocystis ST7-B. Dose–response curves for puromycin, trimethoprim, and WR99210 were estimated from a resazurin-based viability assay (n = 3 independent replicates per drug per concentration). Points show mean ± SD, and the insets list the estimated IC50 values with R<sup>2</sup>-values > 0.75 for all fitted curves. The IC<sub>50</sub> estimates represent the drug concentrations that reduced resazurin-based metabolic activity by 50% under the assay conditions.”

      Right panels: “small-culture complete growth inhibition assay using 1 × 10<sup>7</sup> WT Blastocystis ST7-B cells per culture, assayed in triplicate across a wide range of concentrations. Cultures were incubated for 2 days, and outgrowth was assessed using phenol red acidification of the medium as a culture-level readout, with yellow indicating growth and red indicating no detectable growth. The yellow-to-red transition was used to estimate the concentration required for complete growth inhibition and to guide the subsequent antibiotic selection strategy for Blastocystis ST7-B transformants.”

      “IC<sub>50</sub> and CGI represent distinct assay endpoints: the former measures partial reduction in metabolic activity, whereas the latter identifies the concentration at which no detectable culture outgrowth occurs under the small-culture assay conditions.”

      (6) In Figure 3B, the unexpected UnaG fluorescence pattern could be due to protein sequestration because the protein is mildly toxic to the cell. This should be discussed in addition to the reasons already provided.

      We thank the reviewer for this thoughtful suggestion and agree that protein sequestration or reporter-associated cellular stress represent plausible alternative interpretations of the observed UnaG fluorescence pattern.

      We considered the possibility of UnaG-associated toxicity during interpretation of these data. However, under the conditions tested, we did not observe clear evidence of a substantial toxic effect: UnaG-expressing Blastocystis ST7-B cells could be recovered as stable lines, maintained under antibiotic selection, and propagated through continued culture. We therefore felt that direct attribution of the observed fluorescence pattern to reporter toxicity would currently remain speculative.

      At present, we consider the biochemical properties of the UnaG system itself to provide a more parsimonious explanation for the observed localisation pattern. In particular, unconjugated bilirubin is highly hydrophobic and would be expected to partition preferentially into lipid-rich cellular environments. This interpretation is consistent with the lipid-rich peripheral and intracellular structures previously reported in Blastocystis ST7-B (Liao et al., 2023).

      We have therefore revised the Discussion to acknowledge that the observed UnaG fluorescence pattern may reflect a combination of reporter-specific biochemical behaviour, bilirubin partitioning, local intracellular environment, or possible sequestration phenomena. At the same time, we avoid assigning toxicity as a demonstrated mechanism in the absence of direct measurements of cell fitness, reporter abundance, or bilirubin distribution. Such experiments would be required to evaluate this possibility rigorously.

      Lines 642-648: “Consistent with this, lipid-rich peripheral and intracellular structures have been reported in Blastocystis ST7-B, potentially providing favourable microenvironments for BR partitioning and contributing to the punctate UnaG fluorescence pattern (Liao et al., 2023). An alternative possibility is that the observed signal pattern reflects reporter sequestration or reporter-associated cellular stress. However, because UnaG-expressing lines were recovered, maintained under selection, and propagated through continued culture, toxicity remains a possible but untested explanation rather than a demonstrated mechanism.”

      Minor Comments

      Figure 2: Parts B and C should also show individual datapoints for better reader assessment.

      We agree that inclusion of individual data points improves transparency and interpretability of the underlying data distributions.

      Individual data points have now been overlaid on the boxplots in Figures 2B and 2C.

      Figure 3A: Separate channels (fluorescence, bright-field, merge) should be shown rather than only the merge. The current overlay is difficult to interpret, especially for colour-blind readers.

      We appreciate the reviewer’s concern regarding accessibility and interpretability. We considered separating the fluorescence, bright-field, and merged channels for Figure 3A. However, this panel was intended primarily as an overview demonstrating reporter detectability within the bicistronic construct context, while the detailed fluorescence distribution is explored more extensively in the subsequent UnaG panels. We therefore retained the merged presentation for Figure 3A. Importantly, the image is not dependent on red–green discrimination, as it combines a greyscale bright-field background with a high-contrast green/cyan fluorescence signal that remains distinguishable through brightness and contrast differences. In addition, colour-blind-friendly lookup tables (LUTs) were used throughout the revised figure set.

      To further improve accessibility, the original red annotation arrow has been replaced with a colour-blind-friendly annotation colour.

      Briefly define system components (P2A, UnaG, smURFP, SNAP-tag) and add an abbreviation list.

      We agree that brief contextual definitions improve accessibility for readers less familiar with these reporter systems. Rather than adding a separate abbreviation list, we have added short explanatory descriptions at the points where these components are first introduced in the manuscript.

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.” Line 515–516: “UnaG, a bilirubin-binding fluorescent protein originally isolated from the muscle of the Japanese eel (Kumagai et al., 2013)…” Line 527: “smURFP (small ultra-red fluorescent protein)…”

      Abstract: “among the most prevalent microbial eukaryote” should be “eukaryotes”.

      Corrected in revised manuscript.

      Conclusion (2nd sentence): unclear what “endogenous regulatory part discovery” means.

      We agree that this phrase required clarification. The intended meaning was the identification and benchmarking of native Blastocystis ST7-B promoter and terminator elements for construct design and toolkit development. We have clarified this directly in the revised Conclusion section.

      Lines 682–683 revised to: “By bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements, …”

      Author contributions: “critical advise” should be “advice”.

      Corrected in revised manuscript.

      Again, we thank the reviewers for their careful evaluation, constructive criticism, and thoughtful feedback on the manuscript. The review process has substantially strengthened the manuscript by helping us clarify the distinction between what is directly demonstrated experimentally and what remains mechanistically unresolved.

      The central methodological conclusions of the study remain unchanged: the toolkit enables selectable transgene expression, recovery of colony-derived lines, and propagation of reporter-positive transgenic Blastocystis ST7-B lines, extending genetic accessibility in this organism substantially beyond the previous transient transfection framework.

      At the same time, the revised manuscript now more explicitly acknowledges important unresolved mechanistic questions, including vector topology, P2A-mediated protein separation efficiency, and persistence in the absence of selection. These are now discussed transparently together with the future experimental approaches that will be required to address them directly.

      We believe the revised manuscript now presents a clearer, more rigorous, and more accessible description of a practical genetic toolkit for Blastocystis ST7-B and hope that the revisions and clarifications satisfactorily address the reviewers’ concerns.

    1. eLife Assessment

      This important study explores whether complex structures that are lost during evolution can re-evolve, which is a long-standing debate in evolutionary and developmental biology. The authors demonstrate that re-evolution can occur if the gene regulatory network that underlies the development of complex traits is maintained. The evidence supporting its conclusions is convincing and the work will be of interest to those studying the evolution and development of complex traits.

    2. Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype re-evolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      Comments on revised version.

      The authors have addressed the concerns in the revision. No further comments.

    3. Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species, and in different ant castes within these species. The Authors show that this is linked to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve as the key genes have earlier developmental roles, such that CRISPr knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      As the authors note in their response this is very difficult to achieve. While the addition of this data would raise this manuscript to an outstanding one, I think the data presented is solid, well-presented and provides novel insight.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Vasquez-Correa and colleagues describes the expression pattern of the ocelli (simple eye) gene regulatory network in ants. They correlate the expression pattern of these genes with the presence and absence of ocelli in different classes and species of ants. The presence of ocelli is a polyphenic trait in ants - understanding the molecular and developmental underpinnings of polyphenic traits is of significant interest to evolutionary biologists, developmental biologists, and ecologists. The authors propose that the presence of the latent expression of the ocellar network in classes of ants that do not display ocelli in the adults may underlie the re-evolution of ocelli within the ant lineage.

      Strengths:

      The strengths of the manuscript are that it is well written, the images are of the highest quality, and the data support the conclusions of the authors.

      We thank Reviewer 1 for their positive comments.

      Weaknesses:

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants. A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are, in fact, responsible for ocelli development in the ant.

      The reproductive caste in ants is typically composed of both winged males and winged queens. We agree with Reviewer 1 that the queen caste, which develop 3 fully functional ocelli, is an important point of comparison in our study to the wingless minor workers and soldiers. Unfortunately, however, laboratory colonies rarely produce reproductive queens, and in the field, queen production in colonies of C. floridanus occurs within a narrow seasonal window, making the collection of queen larvae particularly challenging for developmental work. In contrast, the winged males, which also develop 3 functional ocelli like the queens for help during mating flights, can be readily generated in the lab throughout the year. Therefore, we use males as a proxy for characterizing ocelli development and GRN in queens and the winged reproductive caste as a whole. Given the deeply conserved gene regulatory networks underlying this trait across insects, we believe this is a reasonable assumption.

      We also agree with Reviewer 1 that using RNAi to knock down genes in the ocelli GRN would improve the study. For completeness of the scientific record, we would like reviewers and readers to know that we actually did, in fact, invest significant effort trying to knock down otd-1 (ortholog of the Drosophila otd gene), which functions as key upstream regulator of ocellar development. In Drosophila, RNAi knockdown of otd disrupts the development of all three ocelli as well as fine morphological features on the anterior of the head. In C. floridanus, otd -1 is expressed in the head capsule and brain (see Author response image 1 in this response). Injection of dsRNA of otd-1into whole soldier-destined larvae, significantly reduced otd -1 expression in the brain relative to its control, while in the head capsule, otd -1 expression remained largely unchanged relative to its control (see Author response image 1 in this response). This indicates that in the same individual, the injected otd -1 dsRNA was able to penetrate and significantly reduce otd -1 expression in the brain, but, was unable to penetrate the head capsule, where otd -1 expression remained largely unchanged. No ocellar phenotypes could be observed in pupae or adults. Therefore, for technical (not biological) reasons, we were unable to knockdown genes in the ocelli GRN in the head capsule. We hope to solve this technical problem in the coming years to add a mechanistic explanation for the latent expression and maintenance of the ocelli GRN in workers that completely lack ocelli as adults.

      Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype revolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      We thank Reviewer 2 for their positive comments.

      Weaknesses:

      Although the molecular pathway is conserved, the mechanism underlying the lack of ocelli in workers remains unclear. In C. floridanus, it could be explained by the evidence of no expression of certain developmental genes, but in other species, e.g. Polyrachis rastellata, is their expression intact, or reduced? There is no control male.

      In addition, HCR in species with partial re-evolution (if their genomes have been sequenced) would be useful to understand the mechanism. For example, there might be differential spatial expression between medial and lateral ocelli.

      We agree with Reviewer 3 that investigating the mechanisms underlying the lack of specific ocelli in these and other species is the next step for this research. Here, our main focus was instead on trying to explain the mechanisms underlying partial reversion of ocelli through the persistence of ocelli GRN expression in adult workers lacking ocelli. We therefore focused on the latent expression of the ocelli GRN in Polyrachis rastellata, a species that completely lack ocelli in adult workers, and how it may have facilitated the partial reversion of a single ocellus in its congener Polyrachis bihamata. Therefore, although we did not reveal specific interruption points in the ocelli GRN in Polyrachis rastellata, our results showing that this species expresses three genes of the ocelli GRN, offers sufficient evidence that this network is conserved and likely facilitated the partial reversion to a single ocellus in P. bihamata.

      We also agree with Reviewer 3 regarding the male control in P. rastellata and obtaining the species in our study that have undergone partial re-evolution. Unfortunately, these ants occur in Southeast Asia and are very difficult to collect. For males in P. rastellata, our colony died before we could try to induce male development. However, given the deep conservation of the network in the males of a genus within the same subfamily (Camponotini), we feel it is reasonable to assume that the network would also be conserved in the males of P. rastellata, especially since the genes we sampled are conserved in workers that do not develop ocelli as adults. As am sure the Reviewer may know that this is a continual challenge of working with emerging models in evodevo.

      Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species and in different ant castes within these species. The authors show that this is linked to to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstatrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      We thank Reviewer 3 for their positive comments.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve, as the key genes have earlier developmental roles, such that CRISPR knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      We agree with Reviewer 3 that functional experiments in species where workers both have ocelli present and absent is a key next step in this research. Please see our response to Reviewer 1 on our failed attempts to achieve this. We are therefore grateful to Reviewer 3 for acknowledging the challenges in trying to establish RNAi and CRISPR in the head capsules of developing workers in these ants. Also, please see our response to Reviewer 2 on the difficulty of finding and collecting these ants, which occur mainly in Southeast Asia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants.

      A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are in fact responsible for ocelli development in the ant.

      Please see our response to Reviewer 1 above.

      Reviewer #2 (Recommendations for the authors):

      For the questions below, if there is no experimental evidence, consider addressing them in the Discussion.

      Do sizes of ocelli different between castes? For example, even workers have 1-3 ocelli, their sizes are smaller than those of males/queens, especially in workers with 1 ocellus. If so, might it be continuous (not binary) changes in downstream gene expression that control ocellus size, with no ocellus below threshold? Does this favor the hypothesis of threshold but not switch?

      We thank Reviewer 3 for highlighting an important point about the size and development of ocelli. Observations suggest that ocelli tend to be larger in queens and males than in workers and soldiers in species with ocelli. However, we lack quantitative data to test this conclusively. We now include a sentence on the Discussion stating that an important avenue of future work should investigate whether threshold or switch mechanisms influencing the presence/absence, as well as size, of ocelli between queens and workers.

      For the species whose workers have a single ocellus, are there variations, e.g. spanning from 0, 1 to 2? If 2, always one medial plus one of the two laterals? If always one, it would be a good control for staining to see up- vs down-regulation of downstream gene expression within the same individual.

      We agree with Reviewer 2 that this is a fascinating approach to our question. We have not observed natural wild-type variation in the number of developing ocelli in the same-sized individuals in the worker caste. However, in a distantly related leaf-cutting ant species (Atta cephalotes) belonging to different subfamily (the Myrmicinae) individuals with different head-to-body scaling within the same colony can vary in the number of ocelli. For example, soldiers of Atta cephalotes include individuals developing one, two, or three ocelli. These configurations can appear as only the median ocellus, only the two lateral ocelli, or even the median plus a single lateral ocellus. Interestingly, these correlations vary with changes in the size and head-to body scaling, suggesting that each ocellus can undergo different degrees of development, with one or more remaining vestigial or completely absent. On the other hand, workers in other species consistently develop a single ocellus, like in workers of Polyrachis bihamata, with no correlation to size or head-to-body scaling. These cases highlight how evolutionarily labile this trait is among workers of different ant species, which supports our proposal that the underlying gene regulatory network remains latent, thereby facilitating the emergence of novel trait combinations. We therefore agree on the importance of comparing the developmental mechanisms underlying these patterns temporally across larval stages and between individuals within a colony. We have now incorporated 2 sentences into the discussion, stating that this will be an important avenue for future work.

      Is there any function of a single ocellus in workers, or just a consequence of incomplete down-regulation of gene expression?

      Thank you again for highlighting these important points that help us to elaborate on the discussion of our study. The functional role of ocelli in species that develop these structures remains largely understudied. However, for some species particularly within the Formicinae clade the function of the three ocelli in workers has been investigated, revealing that they serve as a celestial compass that facilitates navigation. We reference these findings in our Introduction and Discussion to illustrate that the presence of three ocelli in workers can represent an adaptive trait. In contrast, the functional significance of a single ocellus or of partially developed ocelli remains an important question. This knowledge gap presents a promising avenue for future research to understand the adaptive value of reduced, partially suppressed ocellar development. We have now added a sentence in the discussion stating this.

      In previous studies, JH treatment can increase the number of ocelli in workers, consistent with its role in promoting reproductive development. In the ocellus developmental pathway, what causes the reduction of downstream gene expression in C. floridanus? Does JH directly regulate their expression?

      We thank Reviewer 2 for proposing yet another interesting question for future investigation, which we have added to the Discussion.

      The only current evidence available in C. floridanus is a recent study (MacMillan et al. 2025), in which minor workers and soldiers were treated with JH at different developmental stages. Unfortunately, no evidence of ocelli induction was observed in JH-treated individuals, suggesting that the mechanisms of ocelli development in C. floridanus might be highly canalized, especially in species that exhibit worker polymorphism (inter-individual variation in size and head-to-body scaling within the worker cate). However, more studies are required to understand why in Monomorium pharonis (no worker polymorphism) ocelli development can be readily induced by JH, while in another C. floridanus (with worker polymorphism) it appears quite difficult.

      "In D. melanogaster, the head develops from the eye-antenna disc" This statement is not correct. The brain does not belong to the eye-antennal disc.

      We thank Reviewer 2 for catching the misspelling. We have changed the name to eye-antenna disc in the sentence.

      Reviewer #3 (Recommendations for the authors):

      It is hard to see the developing ocelli in Figure 7 - could the authors increase the contrast to make them more visible?

      We have made the suggested changes to Figure 7 in the main article, and it has indeed improved the figure.

      Author response image 1.

      RNAi knockdowns in developing soldiers of Camponotus floridanus show a reduction of otd -1 expression in the brain, but no effect on otd -1 expression in the eye-antenna disc. A. HCR revealing otd -1 expression in the brain B. qPCR of otd -1 expression after RNAi knockdown shows significantly reduced otd -1 expression in the brain, C. HCR revealing otd -1 expression in the eye-antenna disc D. qPCR of otd -1 expression after RNAi knockdown shows no significant affect on otd -1 expression in the eye-antenna disc.

    1. eLife Assessment

      Building on earlier studies, this manuscript reports a role for pol kappa in cisplatin resistance in the very specific scenarios of head and neck squamous cell carcinoma, providing evidence that the PIP box of Pol kappa is critical for cisplatin resistance in these cells. The findings are of a highly focused relevance and will be useful in the field, but the conclusions are limited to very specific cancer cells. Conclusions cannot be generalized to all cisplatin resistance mechanisms and cell types and are based on incomplete evidence that presents uncertainties and discrepancies that need to be resolved.

    2. Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA-interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data. USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism. The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Overinterpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polk in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014). Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

    3. Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

    4. Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Thank you so much for appreciating our efforts to demonstrate role of Polκ mediated axes in cisplatin resistance in head and neck cancer cells.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      We completely agree with you, and the presented data only support Polκ's role in HNSCC as demonstrated in both acute cisplatin exposure as well as the cisplatin-resistant HNSCC models. Other cell and cancer types may adopt different strategies for cisplatin resistance.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      Since H357-S and SCC9-S cells are highly sensitive to cisplatin, knocking down of Polκ unlikely will alter the phenotype, as other TLS DNA polymerases like Polκ and Polκ play critical role in such lesion bypass. Since no other DNA polymerase was upregulated in these cells upon cisplatin exposure and in the cisplatin-resistant cells, it was intriguing to demonstrate a direct role of Polκ in chemoresistance and that has been proven in this study. Since the catalytic activity of Polκ is not required to induce chemoresistant in these cells, we strongly believe that Polκ-dependent mutagenesis play minimal or no role in adapting cells to tolerate cisplatin. Nevertheless, we will knock down Polκ in these cells and determine cisplatin sensitivity

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data.

      We appreciate the Reviewer's critical and insightful comment. In our view, the interaction between Polκ and USP18 is very specific as USP2 and IgG alone do not pull down Polκ. Similarly, we also show that both Polκ and USP18 interact with PCNA. We agree with the reviewer that the existence of a stable complex of PCNA-Polκ-USP18 has not been fully demonstrated in the current version. We will perform additional experiments to strengthen our finding: a) Co-IP experiments with Polκ PIP mutants (wild-type vs. mutant) should be performed to determine whether USP18 loses its ability to bind PCNA in the absence of Polκ-PCNA interaction. b) Mapping the domain in Polκ that is involved in USP18 binding and their Co-IP experiment. Additionally, high resolution co-localisation images including quantified data will be provided.

      USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism.

      We find the point raised by the reviewer is very intriguing, however, as it will require a significant amount of time and effort to demonstrate ISGylation of DDR proteins and deISGylation by UPS18, and the insight that we may gain is unlikely to add to the central theme of this paper, we will expand this in our subsequent related study. Thank you for the suggestion.

      The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Over interpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polκ in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014).

      Thank you very much for the suggestions. We will take care of the portions and modify as suggested. The necessary reference will be added as appropriate.

      Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

      Thank you for pointing this out and we will modify the portion for better clarity.

      Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Thank you for appreciating our finding that the PIP box of Polκ is critical for cisplatin response

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      We are extremely sorry for the lack of clarity. Please note that two cisplatin-resistant models (H357 and SSC9) have been used and the results were very consistent in both cells. Fig. 1C clearly demonstrates about the generation of these resistant models and the original reference has been already cited.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

      Thank you for pointing this out. In our view, the increased G2 arrest is due to more fork stalling or collapsed than the new origin firing. Also, in our assay we observed less than 10% of new origin fired DNA fibres, and that could be the reason of no significant change in new origin firing among various cells.

      Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      Thank you very much for finding our study interesting and the pending concerns will be addressed as suggested.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      In this study, we have explored eight different cell types (breast, brain, liver, head and neck, pancreatic, prostrate, lungs, and kidney) to check the expression of Polκ upon cisplatin exposure, and HNSCC cells only showed Polκ up-regulation. Therefore, we went ahead for further demonstration of the role of Polκ in cisplatin resistance in OSCC using four different cell models (H357-S, H357-R, SSC9-S, and SSC9-R). By adding more cell lines to study will unlikely change the central theme of the paper. Yes, by acquiring and analysing clinical samples from the cisplatin responder and non-responders would have strengthen our finding.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      It’s an interesting point and we will check whether overexpression of Polκ in H357-S cells could induce resistance to cisplatin and alters IC<sub>50</sub>. Thank you for the suggestion.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      We appreciate your suggestion. As suggested by Reviewer #1 also, we will check the phenotype of Polκ knockdown H357-S cells.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      Thank you for the suggestion, we will modify the text accordingly.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      Please note that GFP-Polκ has been overexpressed in H357 Polκ knockout cells to nullify the effect of endogenous Polκ, otherwise we will not be able to test the role of various Polκ mutants.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      We are extremely sorry for the lack of clarity. The inference has been derived from two sets of experiments as shown in Fig. 4C and Fig. 4D (and is with HU).

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      It has already been demonstrated in Fig. 9C where by knocking down USP18, the DDR proteins like Chk1, Chk2, CtIP, and Artemis can be recovered for ubiquitin-mediated proteasomal degradation. The same results are also obtained when its interacting partner Polκ is deleted. In our view, the presented results have sufficiently demonstrated the role of Polκ-Usp18 in the repair of cisplatin adducts through DDR proteins.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

      We appreciate your suggestion. Since the Usp18 protein is not readily available, we will not be able to show; however, we believe the interaction is direct, and we will be able to map the binding site in Polκ.

    1. eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

    3. Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and postmortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodents vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

    5. Author response:

      eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

      We thank the Editor and the Reviewers for taking the time to carefully review our manuscript and for providing constructive comments and suggestions, as well as the opportunity to revise our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

      We sincerely thank the reviewer for the positive evaluation of our study and for highlighting the significance of examining synaptic tagging and capture following theta-burst stimulation (TBS) in rodents and non-human primates.

      We agree that TBS is a physiologically relevant induction paradigm and that differences in inhibitory circuit dynamics may also contribute to the species-specific effects observed in our study. As highlighted by Larson and Munkácsy (2015), repeated bursts delivered at theta frequency (~5 Hz) can transiently suppress feed-forward inhibition through GABAB receptor-mediated mechanisms, thereby enhancing postsynaptic depolarization and facilitating NMDA receptor activation. We therefore agree that species differences in inhibitory regulation and burst-evoked depolarization may contribute to the distinct expression of synaptic tagging and capture observed between rats and non-human primates.

      We further agree that analysis of the “depolarization envelope” during sequential bursts may provide additional insight into the relative strengths of early induction mechanisms. We will therefore perform these analyses using the existing recordings and compare the depolarization envelope between rodents and NHPs in the revised manuscript. Following the reviewer’s suggestion, we will expand the Discussion section to acknowledge the potential contribution of inhibitory circuit dynamics and depolarization envelope differences during sequential bursts.

      Importantly, however, we believe that differences in downstream molecular maintenance mechanisms also contribute to these species-specific effects. In support of this, our molecular analyses revealed enhanced recruitment of plasticity-related proteins and transcriptional pathways in NHP hippocampus following TBS, including increased expression of BDNF and PKCζ. These findings suggest that both induction-related network properties and downstream molecular stabilization mechanisms may collectively contribute to the enhanced associative plasticity observed in NHPs.

      We also thank the reviewer for the important point regarding PKMζ antisense experiments and the study by Mei et al. (2011). We agree that the absence of an effect of PKMζ antisense oligodeoxynucleotides does not necessarily exclude a role for persistently elevated PKMζ in the maintenance of theta-burst late-LTP. As demonstrated by Mei et al., BDNF together with theta-burst stimulation can maintain late-LTP in the absence of protein synthesis, potentially through stabilization of PKMζ protein levels by reducing degradation rather than through de novo synthesis. However, these findings are not directly comparable to our study, since our experiments involved theta-burst stimulation alone without exogenous BDNF application. Interestingly, our results suggest species-specific differences in the interaction between BDNF and PKMζ signaling pathways. In rats, TrkB/Fc-mediated blockade of BDNF impaired TBS-LTP maintenance, whereas PKMζ inhibition alone had no significant effect. In contrast, in NHP hippocampal slices, inhibition of either BDNF signaling or PKMζ alone failed to abolish late-LTP, whereas simultaneous inhibition of both pathways disrupted LTP maintenance.

      These findings suggest that endogenous BDNF signaling and PKMζ may operate through partially redundant or compensatory mechanisms, particularly in the primate hippocampus. Therefore, although our findings indicate that de novo PKMζ synthesis may not be strictly required under the present experimental conditions, we cannot fully exclude the possibility that protein synthesis-independent stabilization or maintenance of PKMζ contributes to theta-burst late-LTP maintenance in rodents or NHPs. We will now clarify this point in the revised Discussion section.

      Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      We thank the reviewer for this important and thoughtful comment regarding statistical interpretation and biological replication. We agree that, particularly for electrophysiological experiments where multiple slices may originate from the same animal, the effective sample size for species-level conclusions should be considered at the animal level rather than solely at the slice level.

      In the revised manuscript, we will clearly indicate the number of biological replicates (animals) together with the number of slices contributing to each electrophysiological experiment, as well as the biological replicates used for qPCR and Western blot analyses. We will also clarify whether multiple slices from the same NHP/rat contributed to the same experimental condition. These details will be incorporated into the figures and figure legends wherever appropriate.

      In addition, we will perform animal-level analyses by averaging slice responses within each animal prior to statistical comparison and, where appropriate, apply hierarchical or mixed-effects statistical models to account for the nested structure of slices within animals.

      We acknowledge that the number of non-human primates (NHPs) available for this study was inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with primate electrophysiology and tissue collection. Consequently, achieving sample sizes comparable to rodent studies is often not feasible in NHP research. Nevertheless, to further strengthen the biological robustness of the findings, we are currently in the process of obtaining additional NHP brain samples and plan to repeat key experiments in an additional 3-4 animals. We believe these revisions and additional experiments will substantially strengthen the statistical rigor and overall interpretation of the study.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      We thank the reviewer for this insightful comment regarding the mechanistic interpretation of the STC findings. In the present study, we selected the 30 min interval based on well-established classical STC paradigms in rodents, where this interval reliably falls within the effective tagging and capture window. Using this experimentally validated interval allowed us to directly compare whether TBS is sufficient to support STC in primates versus rats under equivalent experimental conditions. Accordingly, the primary objective of this study was to determine whether TBS-induced STC varies across species, rather than to comprehensively define the temporal dynamics of the tagging window.

      We agree, however, that the current experiments do not distinguish whether the primate-specific effect reflects prolonged tag persistence, enhanced plasticity-related protein (PRP) synthesis, altered capture efficiency, or a shifted temporal window. Addressing these possibilities would indeed require systematic temporal interval analyses (e.g., ±15, ±30, ±60, and ±90 min), which represent important future directions. Such experiments are particularly challenging in non-human primates because the availability of primate tissue and experimental resources for large-scale electrophysiological studies remains limited and is currently beyond our experimental capacity due to substantial ethical, logistical, financial, and technical constraints.

      Nevertheless, we fully agree with the reviewer that these experiments are important for advancing the mechanistic interpretation of the findings. Similar temporal analyses have recently proven informative in our rodent studies (Chong YS, Ang SR, Sajikumar S. Commun Biol. 2025;8:553). Importantly, we are currently in the process of obtaining additional non-human primate samples and plan to extend the present work by examining an additional 60 min temporal interval to further characterize the temporal properties of synaptic tagging and capture in non-human primates.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and post-mortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      We thank the reviewer for raising this important and insightful point. We agree that differences in developmental stage between the experimental groups represent an important consideration when interpreting potential species-dependent effects. In the present study, rat experiments were performed in 5-7 week-old animals, whereas non-human primate (NHP) tissues were obtained from 5-7-year-old monkeys. This difference largely reflects the practical, ethical, and logistical constraints associated with NHP research and tissue availability. We acknowledge that these ages are not developmentally equivalent and that maturation state may influence BDNF signaling, protein synthesis capacity, synaptic plasticity thresholds, and transcriptional responses relevant to late-LTP and STC mechanisms.

      We also recognize that differences in euthanasia procedures, tissue extraction, slice preparation, and postmortem handling between rodent and primate tissues may influence tissue physiology and electrophysiological properties. Although extensive care was taken to optimize tissue viability and maintain stable recordings within each species, these variables cannot be completely excluded as contributing factors to the observed differences.

      Accordingly, we will revise the Discussion section to more explicitly acknowledge these limitations and clarify that our findings support potential species-dependent differences under the present experimental conditions, rather than definitive intrinsic species-specific mechanisms. Nevertheless, despite the inherent challenges associated with NHP electrophysiological studies, we believe that the present findings provide an important initial framework for understanding the translational relevance of synaptic tagging and capture mechanisms across species.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      We thank the reviewer for this important comment regarding statistical reporting and interpretation. We agree that the repeated occurrence of identical exact p-values in several nonparametric analyses reflects the relatively small sample sizes and the discrete nature of the statistical distributions. This issue is particularly relevant for the NHP experiments, where biological replication is inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with obtaining and processing primate tissue.

      In the revised manuscript, we will provide exact n values for all comparisons, including the number of biological replicates (animals) and slices where applicable. We will also include additional statistical details, including effect sizes and confidence intervals where appropriate, to improve transparency and facilitate interpretation of the reported findings. Furthermore, we are currently in the process of obtaining additional NHP samples and will attempt to include more biological replicates in the revised version to further strengthen the robustness of the analyses.

      We also agree that the issue of multiple testing should be addressed more explicitly, particularly because multiple genes and proteins were examined. In the revised manuscript, we will clearly state the statistical correction methods applied for multiple comparisons where appropriate. For analyses in which corrections were not applied, we will provide justification, noting that several experiments were based on hypothesis-driven candidate targets rather than exploratory large-scale screening analyses. These statistical considerations will be clarified in the Methods and Results sections.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

      We thank the reviewer for this thoughtful and balanced assessment of our work. We agree that the present data primarily support the conclusion that, under the specific experimental conditions examined, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than that observed in rat slices. We also agree that broader interpretations regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and potentially redundant BDNF/PKMζ-related mechanisms require additional mechanistic investigation and experimental validation.

      Accordingly, we will moderate these interpretations throughout the revised manuscript and clearly state that these conclusions remain preliminary. We will further emphasize that additional experiments, including increased biological replication, expanded temporal analyses, and further mechanistic investigations, will be necessary to more conclusively define the basis of the observed species-dependent differences. Within our current experimental capacity, we are actively working to obtain additional non-human primate samples and plan to incorporate additional biological replicates and key follow-up experiments in the revised version to further strengthen the robustness of the findings.

      At the same time, we believe the present study provides an important initial contribution to an understudied area by directly examining synaptic tagging and capture mechanisms in the primate hippocampus. Given the limited availability of non-human primate electrophysiological data in the field, these findings may offer a valuable framework for future studies investigating the translational and evolutionary relevance of associative synaptic plasticity mechanisms across species.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodent’s vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

      We thank the reviewer for this thoughtful and insightful comment, as well as for the encouraging appreciation of our long-duration plasticity recordings and associative plasticity experiments, which are both technically demanding and time-intensive. We fully agree that interpretation of cross-species differences in synaptic plasticity requires careful consideration of multiple biological and environmental variables, including circadian state, enrichment conditions, strain differences, diet, lighting conditions, and species-specific behavioral ecology.

      Regarding the specific concern related to circadian phase and sleep-wake state, the reviewer raises an important point. Rats are nocturnal animals, whereas macaques are diurnal, and hippocampal plasticity mechanisms are known to be influenced by circadian rhythms and sleep-dependent regulation of synaptic proteins and signaling pathways. Previous studies have demonstrated modulation of LTP, synaptic tagging and capture and protein synthesis in rats across normal sleep-wake cycles. We therefore agree that these factors may influence plasticity outcomes and should be carefully considered in comparative studies.

      Studies have further shown that theta frequency is highly sensitive to sleep-related manipulations. Specifically, theta frequency decreases immediately after sleep, remains elevated during sleep deprivation, and rapidly declines following recovery sleep. In aged animals, these effects appear comparatively attenuated, suggesting reduced sleep-dependent modulation of theta dynamics with aging. Therefore, disruption of normal circadian or sleep-wake patterns may significantly alter theta activity and associated plasticity mechanisms within a species and may not accurately reflect physiological baseline states (Utku Kaya et al., 2026).

      In our experiments, recordings from rats and macaques were performed during their respective active phases under standardized laboratory housing conditions, and we will further clarify these details in the revised Methods section. Nevertheless, we acknowledge that circadian state and related physiological variables cannot be completely excluded as contributing factors to the observed differences between species.

      More broadly, we agree with the reviewer that the present study does not permit definitive conclusions regarding universal “rodent versus primate” rules of synaptic plasticity. Our intention was not to propose a generalized dichotomy between rodents and primates, but rather to report that, under the experimental conditions used here, SC-CA1 TBS-LTP and associated synaptic tagging mechanisms differed between rats and macaques. We agree that broader evolutionary or cognitive interpretations would require systematic comparative analyses across multiple species, including both nocturnal and diurnal rodents as well as diverse primate species. Such studies would provide a stronger framework for distinguishing conserved versus species-specific mechanisms of plasticity.

      At the same time, we believe the present findings remain important because they provide one of the first direct experimental comparisons of SC-CA1 TBS-LTP-associated plasticity mechanisms between rodents and non-human primates under controlled ex vivo conditions. Although the interpretation should be done cautiously, the observed differences raise the possibility that certain metaplastic or protein synthesis-dependent mechanisms may not be fully conserved across species. Accordingly, we will revise the Discussion section to better emphasize the exploratory and comparative nature of the study, while explicitly acknowledging the limitations and potential confounding factors highlighted by the reviewer.

    1. eLife Assessment

      This important study assessed the replicability of a selection of lab-based biomedical experiments in papers published by authors based in Brazil. The study adds a unique perspective to the literature on replication, and provides rich data on the approach taken, the outcomes, and the challenges involved in conducting large-scale crowd-sourced research. The evidence supporting the claims is convincing, but there is scope for clarifying the presentation of the results and extending the discussion section.

    2. Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      (2) The article appears to oscillate between:

      i) a description of the approach and the inherent challenges of such a multicenter replication program.

      ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections. Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers

    3. Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was under represented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX). In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this. If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications. I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean. An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made. In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

    4. Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      We thank the reviewer for the comments.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      We can try to simplify the description of protocols at specific points for example, by providing an overarching description of the study design in the beginning of the Methods, rather than citing our previous eLife paper (Amaral et al., 2019), as suggested below. The methods are indeed quite extensive, but the this may be inevitable in a large-scale project such as this and we note that Reviewer #2 thought that part of the supplementary material should be incorporated back in the main text, which is a suggestion in the opposite direction. It may thus be hard to strike a balance between readability and comprehensibility that can address both reviewers’ opinions.

      (2) The article appears to oscillate between:

      (i) a description of the approach and the inherent challenges of such a multicenter replication program

      (ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      There is a bit of redundancy between tables and text, but this was intentional to make both of them self-explanatory. We also think stating the results in the text can allow us to make each of the replication criteria clearer, a concern that was also mentioned by the reviewer.

      As for requiring particular expertise in statistics for understanding, we mostly disagree. The main results (Tables 1 and 2, Figure 2) are expressed as percentages, and the only statistical concepts needed for interpreting these results are understanding prediction and confidence intervals. For this, we could provide a bit more guidance on their interpretation in the Methods section. Beyond that, most of the secondary results (e.g. Figure 3 and Figure 4) involve linear correlations, which is about as simple as statistical analysis gets.

      Of the results presented in the main manuscript, only Table 3 contains anything beyond percentages and correlations. We do agree that the meaning of each ratio in this table could be more clearly described, but there are essentially no expert-level statistics involved in their calculations.

      Other than that, the main statistical issues are the ideal way to aggregate the results from different replications for which we use different strategies for robustness purposes. However, all of these results are already in the supplementary material, so we don’t feel they interfere to much with the readability of the main manuscript.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections.

      This is indeed a good idea, and we plan to include an initial overarching description of the project in the Methods section of the revised manuscript.

      Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      Again, this is the opposite of what was suggested by Reviewer #2, so we would rather keep the Methods section more or less at its current level of detail.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      The reviewer is correct in his interpretation. Evaluating the main findings of articles or cleaning a field of wrong statements was never a goal of our study (and we were clear about this from the start). Our aim with the project was metascientific (i.e. evaluate the reproducibility of biomedical experiments with a set of common methods) rather than driven by a particular interest in the findings themselves. This is reflected by our choice of selecting experiments from a random sample of articles from multiple fields, rather than filtering by area of interest or importance. It also underlies our choice to evaluate experiments rather than claims, as this was more statistically tractable and potentially more objective as a meta-research goal.

      To be clear, we don’t feel this approach is inherently better or worse than evaluating claims in the literature, as in the Drosophila immunity article case (i.e. Westlake et al., 2026), which is also an important goal. They are merely approaches that answer different questions. Ultimately, we probably made our choice based on (a) our expertise/interest in meta-research rather than in the fields the replications stemmed from and (b) an attempt to engage Brazilian researchers in the project in a way that was non-confrontational and minimized backlash from their peers. We feel this was valuable for many of the lessons learned, although it also meant learning less about the research findings in question.

      Even though this was not a goal of the study, there is some knowledge obtained about the findings that is indeed largely absent from the current manuscript. We do not feel the current format allows for much discussion of 45 different findings, but we do have plans to address these in future articles (as outlined in our response to point 5). In the meantime, qualitative descriptions of each experiment can be found at https://osf.io/w5z9a. This is already mentioned in the Methods but could be reiterated in the results as well.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      We did not go too deep into that finding because we are publishing a separate article focused on the prediction project, which should look into factors that correlate with prediction accuracy, both at the level of predictors (e.g. research field, career level) and of individual predictions (e.g. information taken into account for each answer). We also feel that, given the multiplicity of predictors in the prediction analyses, these findings are a bit tentative, as the strongest predictors may be subject to effect size inflation from the “winner’s curse” effect (as outlined by Reviewer #2). We can try to emphasize it a little more in the discussion (although it already merits a whole paragraph on pages 23-24), but we feel we would be able to discuss it more critically in a follow-up article.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers.

      Indeed, some of our results did not fit this overarching analysis and were left for future publications. One of them is already available as a preprint, while the others are currently in preparation. Specifically, other results from the project should be spread about across five different articles.

      (a) A narrative article focused on challenges and lessons learned with the project, already published as a preprint at https://osf.io/preprints/metaarxiv/8y3tg_v1 (Amaral et al., 2026).

      (b) An article analyzing the prediction survey and markets results in detail (following the pre-analysis plan detailed in https://osf.io/6av7k/files/pjhgd and adding some exploratory analyses on prediction rationales).

      (c) Three articles describing the results of specific experiments with each experimental method (MTT, PCR, elevated plus maze) along with a discussion of aspects inherent to the method that seem to influence reproducibility.

      We can add this information more explicitly to the Methods section, including the links to the papers that have already been published at the time the manuscript is revised.

      Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was underrepresented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Thanks!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      We once more thank the reviewer for the compliments.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      We do acknowledge that the article currently includes a lot of supplementary material. This includes both supplementary figures/tables relating to the paper and many supplementary methods files (mostly hosted at the Open Science Framework). However, we also note that this is already a rather long paper as it stands and that Reviewer #1 has made the opposite suggestion of simplifying it. Thus, it may be hard to strike a balance that will suit all preferences, and we feel that maybe our attempt has landed somewhere in the middle of both reviewers’ ideal versions of the paper.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      We can try to engage with some of the above-mentioned literature in more depth in particular replication studies from other fields (some of which have appeared after our preprint (e.g. Tyner et al., 2026) and with the risk of bias and transparency literature (e.g. Serghiou et al., 2021). That said, we note once more that the article (and the Discussion section) are already quite long, and that analyzing each of these articles in depth is likely to be unfeasible.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Our original plan was to use p values as a predictor (see protocol at https://osf.io/9rnuj), but we later realized this was inadequate as it did not account for effect direction (i.e. significant effects in the opposite direction as the original may yield low p values, but this should not count as replication success). We thus switched to t values to be able to assign positive and negative signs depending on effect size direction. We note that, as we are using non-parametric Spearman coefficients (in which the module of t correlates negatively with the p value), the two approaches are effectively equivalent when original and replication effects have the same direction. This change was accounted for and justified in our list of protocol deviations at https://osf.io/9hj7t.

      Effect size (in relative terms) is already being used in the second predictor in the analysis (i.e. effect size decrease), as our idea was to use one significance-based predictor and one effect size-based predictor, to match what was done for the replication rates). We feel that using relative effects (e.g. response ratios) by themselves may not be as adequate, as for experimental methods with large coefficients of variation and/or low sample sizes (especially PCR ones), one can find large relative effects that are nevertheless far from statistical significance. This also makes relative effects not very commensurable between methods.

      We do believe there is a fair argument, however, to use standardized effect sizes as an alternative to t values (i.e. difference measured in standard errors of the mean) to measure significance/evidence strength. As some replications ended up underpowered, low t values may sometimes be due to insufficient statistical power/low sample size rather than replication failures. Using standardized effect sizes is not devoid of pitfalls (e.g. they can be quite variable when sample size is low), but it is worth doing as a robustness analysis.

      That said, there are a few statistical issues to be decided on how to calculate this (e.g. whether studies should be meta-analyzed using standardized mean differences rather than relative ones for this purpose, or whether an analog of the standardized effect size should be calculated for the log ratio of means). We would have to look more carefully into the multiple possibilities to decide on the best approach (and we do accept suggestions!).

      In the meantime, we note that running the prediction analysis using only experiments with ≥80% power yields a slightly higher correlation of t scores with researcher predictions (ρ = 0.49, p = 0.005), so we do not think that these underpowered experiments affect the trend too much. If anything, they could be masking a higher correlation between researcher predictions and replicability.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      We agree that we should make the description more precise (e.g. “reaching the same results when analyzing a set of data in the same way” for reproducibility and “finding similar results with new data collected under similar conditions” for replicability). We will update these definitions in the revised manuscript.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      We would argue that the reader would be correct in this case: the argument is a bit speculative. It does go in the direction of what is generally accepted within the field (i.e. that publication pressure can lead to lower reproducibility for a range of factors), but we’re not sure this connection has been demonstrated empirically, except for indirect evidence (such as the lower reproducibility in papers stemming from top institutions and “trophy journals” in, the higher frequency of positive results in US states with more researchers in Fanelli, 2010, or the higher number of problematic images for highly productive researchers in some countries in Fanelli et al., 2022. We could cite this evidence in the introduction and make the speculated connection more explicit, perhaps adding modeling work as well (e.g. Ioannidis, 2005; Smaldino & McElreath, 2016) to explain why this could be the case. But essentially, our opinion is that the connection remains a speculation.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      We can offer a more detailed description of the recruitment process (e.g. number and distribution of lectures, social media strategy used, etc.), although we would rather do this in a supplementary document so as not to make the Methods section even lengthier. We note, however, that we never aimed to recruit a “representative sample” of labs from the country: we were busy enough trying to get enough labs for the project to happen, and aware that the call would be inevitably biased by our own communication capabilities and personal networks.

      That said, the response rates for different regions of Brazil do generally match the distribution of research labs and graduate programs within the country (with some distortions likely caused by our personal networks, such as the large number of labs in Rio de Janeiro state), and seem to indicate a rather wide dissemination of the call. One way to visualize this would be to present the distribution of corresponding articles from the original studies selected for the replication (or even from the whole sample of articles obtained for experimental selection) along with the distribution of labs at different stages of the project in Figure S3, which generally show similar patterns. This would actually lend support to our statement that “the population of labs that performed replications was largely similar to the one that produced the original results” in the discussion.

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      We tried to find the combination of methods that would maximize the number of labs that would be included in the project. This is explicitly stated in our Methods Selection document at https://osf.io/qxdjt, but could be stated more explicitly in the paper as well.

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      We initially used single screening by three different reviewers (see https://osf.io/6av7k/files/u5zdq for criteria), as we were merely looking for a sample of experiments; thus, comprehensive inclusion of all eligible studies was not a priority. After this initial screening step, inclusions were confirmed in a consensus meeting with the three reviewers involved.

      Data extraction was also done by a single individual, but the resulting data led to a protocol that was later checked by two reviewers who had access to the paper and were explicitly oriented to judge whether the protocol consisted in a valid replication. Thus, discrepancies between what was in the paper and what was included in the protocol could potentially be flagged at these stages (as they were in many cases). We do note, however, that this is likely not as effective to prevent errors as having data extracted independently, as reviewers may overlook mistakes more easily when comparing two documents rather than extracting data anew. We did find that some errors in extraction slipped by, such as an MTT experiment where treatment concentration was inadvertently changed from mM to μM in a particular protocol step; this was picked up and corrected by 2 out of the 3 labs, but not by the third one, leading the latter replication to be invalidated.

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      As stated in the manuscript, we initially capped experiments at a predicted cost of R$ 5.000 (around USD 1336 at that time), considering reagent cost alone (as equipment and labor was provided by labs), as mentioned in the manuscript. Exclusion rates for that reason were 12/74 (16%) for MTT experiments, 36/132 (27%) for PCR ones and 4/40 (10%) for EPM ones. This is stated at

      This turned out to be an underestimation in many cases, especially as it did not account for pilot experiments, need for repetition, etc; thus, many experiments ended up costing considerably more than that ceiling. As we had included a contingency fund for those cases which we expected would occur , we avoided removing experiments from the sample for this reason as much as possible. Nevertheless, one elevated plus maze experiment ended up not being replicated for cost reasons, as the necessary rat strain was provided by a single facility in the country, meaning that a large number of rats would have to be acquired and transported to all labs at a cost that we were not able to cover.

      As these costs were covered by the coordinating team, we do not feel that this is likely to underlie the reduction in geographical coverage. Other reasons related to lab structure could have led to labs in less well-resourced regions to leave the project, but they probably has nothing to do with the experiments selected.

      That said, the cost cap does mean that the selection of experiments is not completely representative of the literature, but is enriched in relatively cheap and simple experiments which were able to perform (which was our next step for selecting the final sample of experiments. Exclusion rates due to lack of lab expertise and/or infrastructure to perform the experiment were 21/56 (37%) for MTT experiments, 67/89 (75%) for PCR ones and 7/34 (21%) for EPM experiments.

      We will try adding some of this information to the flowchart in Figure 1, as we agree it provides more context on the representativeness of the selected experiments.

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      The scale ranged from 1 (No relevant differences) to 5 (Very relevant differences that prevent considering the study as a direct replication). This scale was used for both the lab and the validation committee scores, and is described at https://osf.io/xgth2 (debriefing protocol) and https://osf.io/e3fjg (validation protocol).

      For the validation committee, we did use a threshold (any score of 4 or a sum of scores of 10 or more among 3 evaluators) to decide what had to be discussed to decide on inclusion, as mentioned on Page 7 of the Methods. For the labs, we used no threshold labs answered the protocol deviation question as a scale, but the decision of whether to consider the study a valid replication or not was not tied to this score.

      We can make both of these points (meaning of the scale and connection to lab’s decision to consider the replication valid) clearer in the Methods section.

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      For the initial analysis of justifications, one reviewer read all answers and flagged those that seemed to concern reproducibility of the methods (e.g. “we replicated the protocol exactly as planned”) rather than results reproducibility (e.g. “effects went in the opposite direction”). We then revised these answers among the whole coordinating team to decide whether we should contact the lab asking them to revise them. We can add this information to the Methods section.

      For classifications of the justification into categories (i.e. Table S7), justifications were classified by two independent reviewers based on categories created after an initial inspection of the data, and discrepancies were resolved by consensus. We can add this information to the table legend.

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      After we extracted data from the lab spreadsheets and summarized the results by code, labs received the results by e-mail and were asked to fill in a form on whether the results were in agreement with what they had found (see details at https://osf.io/nfr6y). Discrepancies in results at least 1 experiment were noted by 36% of the 53 (out of 56) labs that responded. Many of these stemmed from the coordinating team misunderstanding issues such as group identity or experimental unit identification in the spreadsheet. Others had to do with different ways to perform calculations (e.g. relative gene expression or % time spent in open arms). In some cases, simple errors in data transcription or typos caused the discrepancy.

      We were also surprised (and concerned) by the number of experiments in which we later found data errors that were not detected by this process (e.g. 18% of total). Our best understanding of this is that not every lab checked the results with the necessary care, as some errors were quite obvious, as in experiments in which sample size was different, or in which group labels were reversed. Ultimately, agreeing with a form that says “did you find any discrepancies?” may have been performed as a box-ticking exercise with little attention, and was probably not the ideal way to check data which led us to start reviewing results in live meetings afterwards. This is discussed in more detail in our challenges article (Amaral et al., 2026)

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX).

      R 4.5.1 was used for the analysis. We can add this information (which was present in the data repository in the R session info.txt file) and provide the R reference in the manuscript as well.

      In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this.

      If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Yes, we did use the escalc() function for this calculation (for both the replications and the original effect sizes). We can mention this in the manuscript.

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications.

      We calculated the coefficients of variation as the pooled SD divided by the mean of both group means. The reviewer is correct about the possibility of small-sample effects in this case (which we were not aware of). We will thus look into the possibility of implementing this via the escalc () function in the analysis of the revised manuscript.

      We also acknowledge that this could be a source of bias in the comparisons between original and replication CVs (albeit likely a minor one). That said, we note that sample sizes are not always larger in the replication for some experiments with large original effects, power calculations sometimes yielded lower sample sizes in the individual replication, albeit infrequently. On average, though, replication sample sizes were indeed larger.

      I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean.

      This is indeed the case; that said, the CV of the original effect also has random error relative to the true population CV and in that case, there is no way to estimate the uncertainty, as we have a single measure of that parameter. So there is probably no way around ignoring uncertainty in this case.

      We also note that we are looking for evidence of systematic CV inflation across all experiments (rather than for a statistically robust comparison between the CVs of any individual replication). For the sake of measuring this systematic inflation, the use of multiple experiments does allow us to estimate variability at the experiment level which should incorporate the lower-level variability between individual replications if this is not included in the model. Thus, we do not feel that our procedure introduced a systematic bias in the analysis at the experiment-level (although one could argue that it may lead to less precision).

      An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      We thank the reviewer for this suggestion, which indeed seems like an option in this case. We will look into this possibility, although we cannot guarantee at the moment that we will implement it, as we were not previously familiar with the method and will have to study it in more detail.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Yes, but there are likely different reasons for that. Labs leaving after being included may have been due to those in less privileged regions of Brazil (e.g. the northern and western regions of Brazil, generally speaking) having more difficulty in persisting in the project. That said, most of the “disappearance” happens between registration and inclusion which usually has to do with the labs not working with the methods that were ultimately included in the project. We also note that most of the states that lose representation were those that had a single lab to begin with, which may make the visual pattern more striking than the actual trend (as states in the South/Southeast also lose labs, but don’t disappear from the map).

      We note again that we never planned to achieve geographical representativeness when recruiting the labs on the contrary, we were aiming to maximize the number of available labs to run the project. That said, we do agree that for the sake of examining whether the population of labs is similar to the one that generated the original experiments (a claim that we do make in the discussion), this representativeness is important to assess. Once more, to allow the reader to evaluate this, we plan to add an additional map to Figure S3 to describe the Brazilian states where the original experiments came from (based on corresponding author affiliations) in which a similar bias towards the South and Southeast Region can be observed.

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      We agree that this would be useful information, and can experiment with the possibility, but our feeling is that the figure will likely become too noisy in cases where the 95% CIs overlap (which are quite frequent). If this is indeed the case, an option to allow the reader to examine this would be better to add an explicit link to the forest plots for each individual experiment (https://osf.io/sx9gv) in the figure legend.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      We selected the highest correlation values for each continuous outcome (t score and lnRR) and presented these separately in the text. This is a systematic way to perform the selection, but is obviously subject to the “winner’s curse” effect. We agree that adding both metrics for each predictor would be a fair way to keep this in perspective for the reader, but we would have to think about how to do this without sounding too confusing (as results for the two main outcomes are quite different).

      We do note, however, that the outcomes are indeed different and are expected to vary independently in some cases. For the correlation with replication probability predictions, for example, the effects in opposite directions would likely be expected, as larger original effect sizes will likely lead to larger probabilities to be assigned, but also to a higher possibility of effect size decrease. This low correlation between outcomes is probably something that should be pointed out and discussed in the revised manuscript.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made.

      We had given this some thought when writing the manuscript – but ultimately opted not to include confidence intervals for our replication percentages and to use the replication rates as descriptive measures only (as done in other replication studies such as (Errington et al., 2021).

      Even though we aimed for our sample of original experiments to be as systematic as possible, it is ultimately constrained by many factors (the choice of methods, the particular expertise of the labs, etc.) thus, adding confidence intervals represents the uncertainty around the replication rate of a very specific population of experiments, which is not directly comparable to those included in other replication efforts in any case.

      We will reconsider whether we should include confidence intervals for replication rates: although doing this for every replication rate in Table 1 and Table 2 may end up being too much information, it could probably be done at least for the replication rates of the main analysis in the text. We note that calculating confidence intervals for percentages is straightforward, requiring only the numbers that are in the table thus, any reader that wants to estimate uncertainty for those rates should be able to do it easily.

      We will also point out the uncertainty around the percentages mentioned in the discussion when comparing our replication rates with those of other studies, which we agree is an important issue to touch on.

      In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Both correlations in that case are non-parametric (e.g. Spearman’s ρ), so they cannot be directly transformed into Pearson’s r without making assumptions about the distribution (which we would probably avoid doing given the very marked outlier in our own). We can calculate a non-parametric confidence interval for our own correlation coefficient by resampling, but we will have to investigate whether this can be done using the available data from (Errington et al., 2021) (which is probably the case if effect sizes for all experiments have been shared).

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      It is indeed interesting, and seems to confirm an intuition that has long been present in the reproducibility field, but actually has little evidence to support it: if anything, there is evidence in the opposite direction in psychology (Youyou et al., 2023), although they looked at cumulative publication number, while we used number of publications in a fixed interval.

      We can expand a bit further on that finding: that said, we do note that the correlation is relatively weak and has a p value of 0.04. Thus, given the multiplicity of predictors would not be that unlikely to occur by chance, even though it seems intuitive. Thus, even though the relationship seems intuitive, we think it should be considered tentative at best and would refrain from discussing it in too much detail.

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Yes, that is what we mean by “incorrect labeling of error bars” (as can be grasped from the cited references).

      We can perform this regression, which seems relatively straightforward to do. That said, we note that another likely cause for outliers at least for cell line studies would be the use of different (and eventually inadequate) experimental units (e.g. having error bars that represent technical replicates of the same measurement rather than truly independent experiments). We suspect that this may have an even greater effect in terms of causing error bars not to express the same thing and the regression will not help in differentiating the two causes.

      We should also note that different types of experiments may be expected to have very different SDs, so the regression is likely to have a lot of error associated with it. In particular, it’s probably worth doing separate regressions for each method, to account for the likely difference in CVs between animal and cell line experiments, for example. This could also help tease apart the two causes above, as the experimental unit problem mentioned above will likely only be observed for cell experiments.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

      Thanks!

      Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      We thank the reviewer for the compliments. Again, a more extensive list of insights can be found in our challenges article (Amaral et al., 2026), which we will cite in the revised version.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

      We acknowledge the missing replications as a weakness, and we hope we have made that point clear in the discussion.

      Concerning the Brazilian research ecosystem, we could try to explore this in more detail in the introduction. In particular, we believe that a better understanding of the Brazilian academic system, including its regional disparities and the general composition of its workforce (which is largely composed of undergraduate and graduate students), can be useful in interpreting some of the findings.

      We can try to provide a bit more context at the end of the introduction (perhaps between the last 2 paragraphs, which would also address a point made by Reviewer #1), and also in different points of the discussion including those comparing replication rates with other studies or discussing infrastructural difficulties, some of which may be specific to the Brazilian context (such as difficulties in acquiring specific reagents or licenses). Still, we reiterate that, due to the lack of studies with comparable samples in other regions, we cannot tease apart the factors that are specific to Brazil from those affecting lab biology as a whole from the data alone.

      References:

      Amaral OB, Neves K, Wasilewska-Sampaio AP, Carneiro CF. 2019. The Brazilian Reproducibility Initiative. eLife 8:e41602. DOI: https://doi.org/10.7554/eLife.41602

      Amaral OB, Valério B, Carneiro CFD, Mota GPS, Neves K, Abreu M, Tan PB. 2026. Challenges for building up confirmatory science in lab biology: lessons learned from the Brazilian Reproducibility Initiative. MetaArXiv, DOI: https://doi.org/10.31222/osf.io/8y3tg_v1

      Errington TM, Mathur M, Soderberg CK, Denis A, Perfito N, Iorns E, Nosek BA. 2021. Investigating the replicability of preclinical cancer biology. eLife 10:e71601. DOI: https://doi.org/10.7554/eLife.71601

      Fanelli D. 2010. Do pressures to publish increase scientists’ bias? An empirical support from US states data. PLoS One 5:e10271. DOI: https://doi.org/10.1371/journal.pone.0010271

      Fanelli D, Schleicher M, Fang FC, Casadevall A, Bik EM. 2022. Do individual and institutional predictors of misconduct vary by country? Results of a matched-control analysis of problematic image duplications. PLoS One 17:e0255334. DOI: https://doi.org/10.1371/journal.pone.0255334

      Ioannidis jpa. 2005. why Most Published Research Findings Are False. PLoS Medicine 2. DOI: https://doi.org/10.1371/journal.pmed.0020124

      Serghiou S, Contopoulos-Ioannidis DG, Boyack KW, Riedel N, Wallach JD, Ioannidis JPA. 2021. Assessment of transparency indicators across the biomedical literature: How open is open? PLOS Biology 19:e3001107. DOI: https://doi.org/10.1371/journal.pbio.3001107

      Smaldino PE, McElreath R. 2016. The natural selection of bad science. R Soc Open Sci 3:160384. DOI: https://doi.org/10.1098/rsos.160384, PMID: 27703703

      Tyner AH, Abatayo AL, Daley M, Field S, Fox N, Haber NA, Hahn KM, Struhl MK, Mawhinney B, Miske O, Silverstein P, Soderberg CK, Stankov T, Abbasi A, Aberson CL, Aczel B, Adamkovič M, Albayrak N, Allen PJ, Andreychik M, Awtrey E, Axxe E, Azevedo F, Bader MD, Bago B, Bailey J, Bakker M, Banik G, Banks GC, Baskin E, Batruch A, Beatteay A, Behr SM, Berente N, Berry Z, Białkowski J, Bodroža B, Boeschoten L, Bognar M, Bokhove C, Bonfiglio D, Bouwman R, Brady TF, Braithwaite SR, Briceño Jiménez G, Brick C, Bricka T, Briker R, Brown AN, Brown GDA, van Aert RCM, Caldwell K, Capitan S, Capitán T, Chandler J, Charles T, Chartier CR, Chawdhary R, Cheng KJ, Chopik WJ, Clark B, Colvin VE, Comer CC, Costantini G, Coupé T, Cummins J, Czernatowicz-Kukuczka A, de Leeuw J, Dobolyi D, Druckman JN, Duan J, Dujmović M, Dunleavy DJ, Durkee PK, Emery C, Esterling KM, Evans TR, Fedor A, Fernández-Castilla B, Fiala N, Field JG, Fong N, Fonseca MA, Freeman ALJ, Freese J, Geiger SJ, Geng J, Getz LM, Geven LM, Gleibs IH, Gonzales DP, Gooty J, Gourdon-Kanhukamwe A, Greculescu C, Griffin SM, Grigoryan L, Grunow M, Gunby N, Hall B, Hanel PHP, Hannon EE, Harper S, Held MJ, Hickman L, Higgins NC, Hippel S, Hoeppner S, Hong S, Hostler TJ, Inzlicht M, Izydorczak K, Jaeger B, Jankowsky K, Jarke-Neuert J, Jensen M, Jokić B, Jolles D, Jolly P, Jones AM, Juanchich M, Kačmár P, Kapoor H, Keljanovic A, Koirala S, Kołczyńska M, Kouroupaki D, Kühnen U, Landgrave M, Larson MJ, Laulié L, Lawrence ACE, Le Forestier JM, Leahy KE, Lee S, Leslie J, Lewis SC, Limnios C, Lin H, Liu A-C, Lloyd JW, Ludvig EA, Lynott D, MacDonald J, Mallik P, Mallinson DJ, Marinazzo D, Martarelli CS, Matacotta J, McBride A, McHugh C, McMillan G, Méndez E, Metzger M, Michaelides MP, Michalak J, Micheli L, Miller JK, Milyavskaya M, Molden DC, Monjaras AG, Moreau D, Morrow A, Moya C, Mudrik L, Mulder LB, Munt KA, Nandi A, Nason K, Nast C, Nave G, Nax HH, Neubauer F, Nguyen PLL, Nichols AL, Nilsonne G, O’Boyle E, Oettinghaus J, Oh J, Oshana A, Ostermann T, Ostrowski RP, Oyebanjo A, Panczak R, Patrianakos J, Pavez I, Pavlov YG, Persson S, Perugini M, Peters K, Pieters C, Ponizovskiy V, Porter ND, Prenoveau JM, Purić D, Purol MF, Puthillam A, Quinn KA, Ramljak M, Reed WR, Ritchie M, Ritzau M, Roche SP, Rodela R, Röer JP, Ropovik I, Rothschild J, Saal J, Safadi H, Samaha J, Sanchez M, Sankaran S, Santos D, Sargent AC, Sauter M, Schmidt K, Schnabel L, Schroeder AN, Schuetz SW, Schuetze BA, Schulte-Mecklenbeck M, Schütz A, Sevigny EL, Shackleton E, Shafranek RM, Shaki S, Shakya S, Sirota M, Sisco MR, Sitnikov MM, Slevc LR, Smalarz L, Smith CT, Snyder JS, Sommet N, Sonmez F, Spellman BA, Stanulewicz-Buckley N, Stock G, Street CNH, Strømland E, Sundelin T, Syed M, Szabelska A, Szaszi B, Szumowska E, Tagat A, Täuber S, Tay L, Thapa S, Thatcher J, Tsaklakidou D, Tummers L, Turkovich E, Tutor MV, Urbanska K, van ’t Veer AE, van Assen M, van de Ven N, van den Goorbergh R, Vargo EJ, Vaughn LA, Vazire S, Vermeulen JM, Vo DTH, Volkman V, Wagenmakers E-J, Wagner D, Walasek L, Walter F, Warmelink L, Wei L, Weißflog MI, Weller N, Wichman AL, Wilbiks J, Williams JR, Wolfe K, Wort F, Wright R, Wulff JN, Xue X, Yan VX, Yang Y, Yoon S, Žeželj I, Zhang Y, Ziano I, Zogmaister C, Zupan Z, Zwaan RA, Nosek BA, Errington TM. 2026. Investigating the replicability of the social and behavioural sciences. Nature 652:143–150. DOI: https://doi.org/10.1038/s41586-025-10078-y

      Westlake H, David F, Tian Y, Krakovic K, Dolgikh A, Juravlev L, Bournonville TE de, Carboni A, Melcarne C, Shan T, Wang Y, Mu Y, Kotwal A, Pirko N, Boquete JP, Schüpfer F, Rommelaere S, Poidevin M, Liu Z, Kondo S, Ratnaparkhi GS, Chakrabarti S, Liu G, Masson F, Xiaoxue L, Hanson MA, Jiang H, Cara FD, Kurant E, Lemaitre B. 2026. Reproducibility of scientific claims in Drosophila immunity: A retrospective analysis of 400 publications. eLife 15. DOI: https://doi.org/10.7554/eLife.108404.1

      Youyou W, Yang Y, Uzzi B. 2023. A discipline-wide investigation of the replicability of Psychology papers over the past two decades. Proceedings of the National Academy of Sciences 120:e2208863120. DOI: https://doi.org/10.1073/pnas.2208863120

    1. eLife Assessment

      This valuable paper uses a mathematical model applied to a dataset of E coli / ESBL carriage and transmission to infer drivers of drug resistance in France. The strength of support for the study findings is incomplete. While the research question is of importance, and the mathematical model has structural and methodological integrity, numerous issues are noted: insufficient description of the data, lack of included equations and code, definitions of antibiotic use that are not complete, low sensitivity of assays for carriage, technical issues with statistical prior selection and parameter identification, and application of non-regional ECDC surveillance data to France.

    2. Reviewer #1 (Public review):

      Summary:

      The authors used a large dataset evaluating gut carriage of Enterobacterales and ESBL organisms from children aged 6-24 months as the basis for a modeling study to investigate what factors are most important for determining the prevalence of ESBL resistance. The modeling incorporated travel, a simple model of carriage duration (short and long), fitness cost of resistance on transmission and clearance, and antibiotic use. They found that antibiotic use is the primary driver of resistance prevalence, with transmissibility of resistant strains also important for setting the prevalence. Travel, while important when prevalence is very low, plays less of a role in maintaining prevalence once it is established (in keeping with other recent work). They estimated the fitness cost of resistance (terming a reduction of 14% on the rate of transmission and an increase of 23% on the rate of clearance as "low"). While the extent of assumptions and simplifications makes me skeptical of the quantitative conclusions, the qualitative ones seem reasonable and reinforce the long-held principles of the field--reducing antibiotic pressure and interrupting transmission--and highlight the importance of understanding the biological factors that shape the duration of carriage and the likelihood of colonization.

      Strengths:

      This study incorporates many of the factors that might influence the carriage prevalence of ESBL Enterobacterales. This builds on the work led by this group, both in primary data collection and in theory. Overall, it's such a tough problem that I commend the authors for trying to tackle it. The authors take a thoughtful, rigorous approach, acknowledging simplifications and assumptions where they need to, so as to evaluate the various factors shaping ESBL prevalence.

      Weaknesses:

      Part of the reason it's such a tough problem is that we have limited data to structure and parameterize a complex model.

      (1) The data are not sufficiently described.

      The primary data source for this modeling exercise comes from a study of 6-24-month-old children who underwent rectal swabs and evaluation of the carriage prevalence of Enterobacterales, and then whether these Enterobacterales were ESBL; moreover, the study included data on travel and on antibiotic use. Could the authors please direct us to these primary data? Could the authors also justify the parameters in their models from these data--for example, could they please provide the distribution of antibiotic use and the associated timing? Could they also explain why they decided to treat all Enterobacterales as if they were E. coli (line 307)? Is there evidence that all Enterobacterales occupy the same niche and compete with each other?

      (2) The model should be more fully described and the limitations explored/explained.

      - The authors should point to the code and the ODEs.<br /> - I understand the focus on the pediatric population; the authors argue that this is reasonable because ESBL colonization is similar across age groups. But presumably, antibiotic use differs across age groups, and there is colonization pressure from within households.<br /> - The authors only consider resistance to extended-spectrum beta-lactams and use of beta-lactam antibiotics, but ESBL Enterobacterales are often resistant to other antibiotics as well. How much does the use of other antibiotics also select for Enterbacterales that happen to carry ESBL resistance? "One bug/one drug" modeling, as done here, neglects the complexities of the actual patterns of resistance and range of antibiotic use.<br /> - Do the data support the T3 or S3 compartments, which, if I understand correctly, means no exposure to antibiotics can happen during three months after either treatment or travel? What do the data say about the patterns of antibiotic use? I'd imagine that the likelihood of antibiotic use is not homogenous, but instead, there are some who use repeated rounds of antibiotics.<br /> - Why do the authors exclude individuals who used antibiotics in the prior 7 days? What justifies that cutoff? The authors speculate that the impact of excluding these individuals is likely to be minimal; why exclude them, then? Did the authors evaluate the results if they were included?<br /> - What is the basis of "niche differentiation", as described starting on line 221? Why should clearance of one strain be slower when the strain co-occurs in a host with a strain of another type?

    3. Reviewer #2 (Public review):

      Overview:

      This study integrates several datasets into a unified modeling framework that incorporates several mechanisms thought to impact the spread of ESBL-resistant bacterial strains. The model accounts for tradeoffs between persistor and colonizer strains, travel rates, antibiotic treatment and strain clearance, direct competitive interactions, and, most importantly, a series of distinct costs associated with the carriage of ESBL resistance. The resulting 75-compartment model is internally consistent and structurally neutral. However, the parameter estimation is flawed in many ways, compromising the interpretations of the model.

      On the usage of the Swedish infant data set to estimate colonization and persistence:

      First, while other papers have taken similar approaches, the Swedish infant data set is fundamentally inadequate to estimate colonization and persistence rates. This is because very few colonies were typed per sampling event (2 to 6 colonies per event). The original authors themselves argued that strains of indistinguishable morphology would not be able to be differentiated by this method. They also provided data showing that strain identity was not directly related to colony morphology (same strain often displaying distinct morphologies).

      The consequence of this is that strains present in low abundance would be missed with a high likelihood. However, if they were to be stochastically sampled, this would count as a "colonization" event, and if they were missed in subsequent samplings, this would count as a "loss" event. In other words, the statistical methods described conflate within-host dynamics (which might lead to distinct within-host abundances) with between-host dynamics (colonization and loss).

      Beyond this conceptual issue, some technical aspects aren't particularly sound. The mean of the inferred posterior for the lambda and mu parameters are then used to calculate the beta, gamma, d, and epsilon parameters through a linear regression. The more technically correct way of doing this would be to directly infer these parameters from the data and obtain a full posterior for these parameters.

      This highlights another issue: these parameters are passed down to the next statistical model as point estimates, with no associated uncertainty. This artificially inflates the (already low) confidence of the estimates for the cost parameters.

      Finally, when this procedure generated parameters that were inconsistent with their expectations (clearance is too high to explain prevalence in France), they adjusted the parameters by discarding and recalculating their beta parameters to artificially enforce neutrality between their strains and enforce the expected prevalence. This is problematic because beta and gamma were jointly estimated, and there is no particular reason why some of them should be discarded. The more natural interpretation would be that parameters inferred from Swedish infants do not translate well to French adults, which should preclude their usage in this context.

      On the estimation of costs of ESBL resistance:

      The core of the second statistical model is to use prevalence data, travel data, and treatment data in conjunction with the previously inferred colonization and loss parameters to infer the costs of carrying antibiotic resistance. Therefore, the accuracy of this section is contingent on an accurate estimation of the previous parameters. However, these colonization and loss parameters are inherited with no uncertainty (just point estimates are passed down), which, as previously mentioned, generates an artificially precise posterior distribution for the resistance parameters.

      However, the most severe issue with the statistics lies in the choice of priors for the cost parameters. All of them are uniform in a positive range that implies a positive cost. Importantly, the average over a positive range will always be positive; therefore, this method will ALWAYS estimate a positive mean for the costs. Note that the posterior distribution of some cost parameters seems to peak around zero and abruptly decays with no mass to the left of zero. This is caused by the choice of prior. Had delta been allowed to be negative (i.e., antibiotic resistance carried a benefit, having the prior be uniform between -1 and 1), the posterior distribution would likely be much more symmetrical, and the confidence interval would have included 0.

      Restating, because the prior is a continuous function between 0 and 1, it contains infinitely more mass in the region that represents there being a cost (delta>0) than in the region representing no cost (delta=0). This means that it is a mathematical impossibility for this model to infer the absence of a cost.

      Therefore, the main finding of the paper ("We found that resistance is costly") is a mathematical artifact of the prior choice and of the model structure.

    4. Reviewer #3 (Public review):

      Cotto and colleagues integrated data analysis with mathematical modeling to examine extended-spectrum beta-lactamase (ESBL)-producing E. coli in France. While ESBL prevalence has risen globally, it has stabilized at approximately 6-8% across Europe. Established risk factors for ESBL carriage include prior antibiotic exposure and travel to high-prevalence regions, most notably South-East Asia. The dataset incorporated information on ESBL-producing E. coli and travel history in young children, and the model was calibrated to ECDC surveillance data on ESBL across Europe, supplemented by literature-derived parameters on antibiotic use, E. coli biology, and transmission dynamics. The authors report that ESBL-carrying strains exhibit a 14% fitness cost in community transmission relative to susceptible bacteria, yet are cleared 23% less frequently. ESBL carriage was strongly associated with factors that prolong gut colonization. Both antibiotic treatment rates and transmission efficiency were identified as key determinants of community-level ESBL prevalence.

      Strengths:

      The study addresses a clinically and epidemiologically important topic. The integrated modeling approach is methodologically sound and well-suited to disentangling the relative contributions of transmission and antibiotic selection pressure.

      Weaknesses:

      Several concerns regarding the data used in this study warrant consideration. First, model calibration relied on ECDC surveillance data pooled across multiple European countries, several of which have substantially lower antibiotic consumption than France (ECDC ESAC-Net Annual Epidemiological Report, 2024). Given that antibiotic use is a primary driver of ESBL selection, ESBL prevalence is likely to be heterogeneous across these settings. Calibrating to a geographically diverse dataset risks introducing systematic bias into parameter estimates that may not be representative of the French context. The authors should repeat the analysis using France-specific data, or, where this is not feasible, restrict the calibration dataset to countries with comparable antibiotic consumption profiles. Second, the travel exposure data may be insufficient to adequately capture importation dynamics from South-East Asia, as the cohort consisted exclusively of young children, a demographic less likely to travel to high-prevalence regions than older age groups. This may result in an underestimation of travel-associated importation as a contributor to community ESBL prevalence, and the generalizability of these findings to the broader population should be interpreted with caution.

    5. Author Response:

      We thank you for this assessment of our work and the positive assessment of the overall theoretical framework. We can fully answer the concerns, in particular regarding data quality, and will provide detailed answers in the following directions:

      “insufficient description of the data”: We will describe the data as much as possible and will share the data and the analysis code.

      “lack of included equations and code”:  We will share the mathematical equations in the supplementary material, and the full code on an online repository. The reason why our data repository (10.5281/zenodo.18480481) is not yet public is that it cannot be changed after publication. For the review process, we provide a github link to data and code here https://github.com/oliviercotto/eLife_epidR. We will ultimately share the link to the final version of the files on Zenodo.

      “definitions of antibiotic use that are not complete”: We will complete the definition of antibiotic use, which is the use of any antibiotic between 7 days and 3 months before sampling. Children who used any antibiotic 7 days before sampling were not included in the study. The type of antibiotic used is given in supplementary material S1: 93% of the antibiotics prescribed are beta-lactams (amoxicillin, amoxicillin/clavunalate, oral 3rd generation cephalosporins).

      “low sensitivity of assays for carriage”: the carriage study conducted in Sweden is used to get plausible estimates of carriage duration parameters in infants.

      • Strain definition is based mainly on randomly amplified polymorphic DNA (RAPD), not colony morphology. Strains with distinct morphology but the same RAPD profile are considered one strain. Conversely, it was checked that strains of the same timepoint with the same morphology most often had the same RAPD profile.

      • We did check that these data are not much affected by imperfect sampling: observations of a strain ‘disappearing’ from sampling then ‘reappearing’  at later timepoints are rare (14 out of 273 strains). This is why we did not correct these occurrences in the previous version of the analysis. In the revised version, we will add a description of these occurrences and correct them. This correction did not significantly alter the inferred parameters in our preliminary analyses.

      • At a broad level, the fact that E. coli clades vary in their carriage duration is very well established across multiple independent datasets; the precise value of carriage duration difference for “persistent” vs. “transient” that we inferred here (a two-fold difference, supplementary material S3) is actually relatively conservative, in the sense that other studies have detected more important differences. We will create a table summarising available evidence on colonization parameters of E. coli to show that the insights from the Swedish data are qualitatively robust.

      • Yet, we will conduct a range of sensitivity analyses to see how the inferred costs of ESBL resistance vary when varying differences in carriage durations, competition and niche differentiation.

      “technical issues with statistical prior selection and parameter identification”: We disagree there is a “technical issue” with prior selection: The fact that resistance is costly, hence that our priors are left-bounded at 0 for the cost parameters, is a prior expectation based on the observation that resistances do not go to fixation. If resistances only conferred an advantage in treatment, but zero cost, then they would quickly evolve to 100% frequency–contrary to what is observed in virtually all epidemiological studies of resistance. That said, we will relax the definition of these priors to test that the data is also compatible with a strong cost on some traits, and no cost or a “negative cost” on other traits.

      Regarding parameter identification and the specific comments on the inference of colonisation parameters: we re-inferred all colonisation parameters with direct inference assuming specific functional forms. This does not alter much the final colonisation parameters that we then use for our main inference. We will also conduct sensitivity analyses to examine how changing some of the colonisation parameters (carriage duration, competition and niche differentiation)  would alter the main inference.

      “application of non-regional ECDC surveillance data to France”: We will clarify our text, as there is a misunderstanding here: we do not use non-regional ECDC surveillance data for inference. We use ECDC data (i) for illustrative purposes, to show that trends in ESBL in France in this surveillance system are very similar to those observed in France in our focal dataset, thus showing the consistency and representativeness of our data. (ii) to give an overview of the weak and inconsistent association of ESBL with age across Europe, thus supporting the relevance of our approach even if our data concerns infants and children. We will make sure this is clarified in the updated version of the manuscript.

    1. eLife Assessment

      This study presents an important finding regarding the effect of Yoda molecules on PIEZO2 function, challenging the assumption that they selectively activate PIEZO1. The evidence supporting this claim is solid, but several methodological and conceptual issues need to be addressed. Overall, this work will be of broad interest to researchers working with PIEZO channels across various biological scales.

    2. Reviewer #1 (Public review):

      Summary:

      In this work, T. Wijerathne et al. investigated and reported the agonistic effect of Yoda1 and Yoda2 over PIEZO2 function using patch clamp electrophysiology, Ca2+ imaging, and molecular dynamics. They find that Yoda1 sensitizes PIEZO2 to membrane tension, can induce Ca2+ influx, and decreases its inactivation to a lesser degree than it does to PIEZO1 channels. Additionally, their data shows that Yoda2 sensitizes PIEZO2 channels to membrane indentation to a greater extent, but it has a weaker effect on channel inactivation than Yoda1. Interestingly, they report that a mutation in a conserved arginine between PIEZO channels can be used to abolish PIEZO1-mediated Ca2+ flux in response to Yoda molecules. As a whole, the results presented here should be put into perspective against previous and future works involving systems where both PIEZO1 and PIEZO2 might be expressed. This is especially true for works where Yoda1 has been used as a basis for determining the absence of PIEZO2.

      Strengths:

      The authors use multiple techniques to investigate how Yoda molecules affect the three most important biophysical aspects of PIEZO channels that, when changed, result in pathophysiological responses: a) sensitivity to mechanical stimuli, b) Ca2+ entry, and c) channel inactivation. Lastly, they find a specific amino acid/region that could be exploited for drug design and/or development.

      Weaknesses:

      The methods and discussion sections are lacking enough detail to fully evaluate the findings and put them into perspective, respectively.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript challenges the long-standing assumption that Yoda1 and Yoda2 are PIEZO1-selective activators. Using patch-clamp electrophysiology and calcium imaging in HEK293TΔPZ1 cells overexpressing PIEZO2, the authors demonstrate that Yoda1 potentiates PIEZO2 stretch-activated currents to a similar extent as PIEZO1 and slows PIEZO2 poking-current inactivation (albeit with lower efficacy). They further show that the more potent analog Yoda2 affects PIEZO2 at nanomolar concentrations and use mutagenesis and molecular dynamics simulations to propose that Yoda2's benzoic acid group forms a transient salt bridge with R1724 in the putative Yoda binding pocket, explaining its enhanced potency.

      Strengths:

      The authors are established Piezo/biophysics experts; the study is highly important, technically competent, and carries significant implications for the reinterpretation of prior work that used Yoda compounds as PIEZO1-selective probes.

      The core finding that Yoda1 modulates PIEZO2 stretch currents is convincing and important. However, several conceptual, methodological, and presentational issues need to be addressed before acceptance, as detailed below.

      Weaknesses:

      (1) The abstract states that Yoda1 potentiates PIEZO2 "as efficaciously as PIEZO1." This claim is accurate only for stretch currents and single-channel open probability, but the paper itself demonstrates important asymmetries: i) Yoda molecules slow PIEZO2 poking-current inactivation ~2-fold, versus ~5-10 fold for PIEZO1 (Figure 3b and ref #60). ii) Spontaneous Ca²⁺ entry via PIEZO2 requires non-physiological conditions (high extracellular Ca²⁺, hypertonic solutions) that are unlikely to occur in native cells.

      The abstract should be revised to clearly qualify where equivalence holds and where efficacy differences exist. IMO, the current wording risks overcorrecting the historical bias (PIEZO1-only) by going too far in the other direction.

      (2) Related concern: the PIEZO2 Ca²⁺ signal in Figure 2 is only detectable using a Ca²⁺-boosted solution (CBS ie 30 mM Ca²⁺). Physiological extracellular Ca²⁺ and cells normally do not experience sustained hypertonicity at these magnitudes. The authors should explicitly clarify that the practical implication of their findings is primarily for electrophysiological (patch-clamp) experiments and that the Ca²⁺ imaging caveat applies only under amplified conditions. Specifically, the authors should state that in standard Ca²⁺ imaging assays with physiological buffers, PIEZO2 is unlikely to confound Yoda1 results.

      Related point: Can cytochalasin D (CytoD) restore a Yoda1-dependent Ca²⁺ signal in physiological saline? This would help determine whether the weak PIEZO2 response is primarily a membrane tension issue (cytoskeletal tethering) versus intrinsically lower channel expression or permeability. The authors already have tagged PIEZO1/2 constructs and could, in principle, normalize by surface expression.

      (3) The mean inactivation tau values for wild-type PIEZO2 poking currents in both DMSO and Yoda1 conditions (Figure 3b, approximately 15-40 ms range) appear substantially higher than values reported in published literature (typically 5-10 ms; eg, PMID: 20813920). This discrepancy needs to be addressed.

      (4) The authors perform all MD simulations on a truncated PIEZO1 model and justify this choice by noting that the Yoda binding region is highly conserved between homologs. This is a reasonable and defensible starting point given the availability of well-validated PIEZO1 simulation set ups in their lab. A few points are nonetheless worth addressing: While PIEZO2 simulations are not strictly required, the authors are encouraged to briefly discuss whether any long-range structural differences between PIEZO1 and PIEZO2 (outside the binding site itself) could influence Yoda2 binding dynamics, particularly in light of the chimera data showing that PIEZO2 sequence in repeat A abolishes Yoda1 sensitivity. This reviewer still doesn't understand the reason behind this discrepancy despite it being acknowledged in the text.

      Another MD-related comment is that three simulation replicas (which is impressive for such a big system) show markedly different salt bridge occupancy (82.6%, 49.7%, 99.8%; stated in the text). This wide variation suggests incomplete sampling in at least one replica. The authors should provide RMSD plots for ligand and protein backbone to assess convergence and possibly discuss whether the 49.7% replica represents a genuinely distinct binding mode or incomplete equilibration.

      (5) The Discussion proposes that PIEZO2's weaker Ca²⁺ response to Yoda1 could partly reflect lower membrane expression. Since the authors already have fluorescently tagged PIEZO1 and 2 constructs, a simple fluorescence intensity comparison between the two (acknowledging it would reflect total rather than surface expression) could provide at least indirect support for this claim. Alternatively, if such a comparison is not feasible, the authors may consider removing membrane expression from the list of proposed explanations or explicitly acknowledging that this remains unsubstantiated speculation. The max poking currents may somewhat and roughly indicate the level expression difference too, if done exactly side by side.

      (6) The abstract or concluding remarks should highlight that Dooku1 is not PIEZO1-selective in its agonist-like action on PIEZO2, and that Cmpd15/Cmpd64 appear to be better PIEZO1-selective tools. This nuance is buried in the Results section.

      (7) The authors should not cite PMID 31015490. Clearly, any work on MCC13 is confounded by the overwhelming expression of PIEZO1 (PMID: 42084270). Instead, the authors should also cite the literature from others who have clearly recorded stretch currents from PIEZO2 before the cited studies (eg, PMID: 37590348).

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript reports that Yoda1 and Yoda2 agonize PIEZO2 in a manner similar to PIEZO1, increasing open probability and stretch sensitivity, but the mechanism underlying this sensitivity is incomplete. Mutagenesis was shown exclusively in PIEZO1, with no corresponding mutagenesis in PIEZO2, so the proposed mechanism in PIEZO2 is inferred by homology rather than directly tested. All experiments use mouse PIEZO2, and the human ortholog should be used before generalizing the proposed reinterpretation of the field.

      Strengths:

      The pressure-clamp electrophysiology demonstrating a shift in half-activation pressure for PIEZO2 is compelling evidence in support of the central claim.

      Weaknesses:

      (1) In the single-channel recordings (Figure 1a), it's unclear how many channels were present in those patches. After applying -60 mmHg pressure, multiple channels would be activated (as seen in Figure 1e). The number of channels in the patch and their inactivation rate could significantly influence the open probability in such experiments. To overcome this, in the original Yoda1 article (Syeda, Ruhma, et al. eLife 2015), no additional pressure was used. Additionally, the reported open probability comparison (n=7 Yoda1 vs n=17 DMSO patches) has an SEM nearly as large as the effect itself (0.30 {plus minus} 0.11), consistent with a small number of outliers driving this. The underlying mean open and shut times are reported without any statistical test; only the derived open probability receives a p-value. Additionally, in Figure 1a, the Yoda1 condition noise is different from the control. This should be stated if noise filtering was applied and how, given that this could affect open probability analysis.

      (2) The calcium imaging data in Figure 2 raise significant concerns regarding the chemical activation claim. The calcium-boosted solution (30 mM Ca2+) is not physiological and appears to be generally stressing cells rather than specifically activating PIEZO2: the control condition under CBS already shows an elevated signal, consistent with cells being unwell at this calcium concentration, and adding Yoda1 on top of this shifted baseline raises further questions about specificity rather than confirming it. Separately, it is unclear why DMSO alone produces measurable PIEZO2-associated calcium influx in HBSS, a result that is not addressed in the text. Figure 2 should clearly indicate when DMSO/Yoda1 perfusion was initiated, and y-axis labels are missing from panels A and B.

      (3) In the poke experiments, an activation threshold should be calculated and reported, and amplitude data (e.g., peak current versus indentation depth) should be shown rather than only inactivation tau values. It is also unclear why mClover3- and N-GFP-tagged constructs were used in these experiments, since electrophysiological recording already confirms channel expression without requiring a fluorescent tag.

      (4) For inactivation kinetics (Figure 3b), the authors use unpaired comparisons across separate cells, whereas the deactivation experiments (Figure 3c) use paired; it should be applied to the inactivation experiments as well. Deactivation kinetics for PIEZO2 itself should be shown. If the claim is that Yoda1 acts on PIEZO2 through the same mechanism proposed for PIEZO1, then a PIEZO1/2 chimera should be expected to show a corresponding effect on deactivation tau; instead, this chimera is reported as completely Yoda1-insensitive despite both parental channels being Yoda1-sensitive, as shown in this study.

      (5) Given that this reflects a different experimental paradigm for Yoda EC50, PIEZO1 should be included within Figure 4b. Additionally, EC50 bar plots should be present on this figure. The inactivation time constant for PIEZO2 without Yoda1 is inconsistent across figures, below 20 ms in Figure 3b but above 20 ms in Figure 4c.

      (6) Finally, the modeling is performed exclusively on PIEZO1, whereas the manuscript's central focus is PIEZO2. It is therefore unclear whether the proposed structural mechanism, including the basis for Yoda2's reduced efficacy on PIEZO2, can be directly extrapolated to PIEZO2.

    1. eLife Assessment

      In this manuscript, the authors describe a new member of the KCNE auxiliary subunits of potassium channels from a lamprey. This new subunit represents an early evolutionary member which confers new properties when expressed along with KCNQ channels. The authors present convincing evidence from several experimental approaches. The contents of this manuscript are important and should be relevant to understanding both the mechanism of modulation of KCNQ channels by KCNE subunits and the evolutionary history of these subunits, which this manuscript now extends to the divergence of early vertebrates.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

      Weaknesses:

      (1) No error bars visible for lamprey Q1 isoforms (open symbols) in Figure 2G. No statistical comparison was provided to indicate whether lamprey Q1 isoform V1/2s are significantly different (nor in Supplementary Table 1).

      (2) There is the same issue in Figures 3 and 4. No appropriate statistical comparison is made between V1/2s for different truncations of PmKCNE0 (Figure 3), or between KCNQ1 species isoforms with and without PmE0.

    3. Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

    1. eLife Assessment

      In this important study, Boudjema et al. use cell culture models and high quality advanced microscopic imaging to provide detailed analyses of the cellular processes underlying centriole amplification, apical migration, and assembly of hundreds of motile cilia in multi-ciliated cells. The authors present convincing evidence showing that in these cells all the molecular and cellular steps controlling centriole biogenesis that in cycling cells extend over almost two cell cycles, occur within a single cell cycle variant. This work provides a better understanding of the regulation and order of these processes and is of interest to all cell biologists and in particular researchers studying centrioles and cilia.

    2. Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      Comments on revised version.

      The authors have significantly improved the manuscript, by refocusing it, introducing text and figure changes, and by adding new data including functional analyses. The revised version now has convincing data that support the claims. All my remaining concerns have been addressed.

    3. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the Cen2-GFP; mRuby-Deup1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described Cen2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of Deup1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells(treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules and microtubule-based transport in different stages of differentiation in brain MCCs. Addressing the role of microtubules during different stages of centriole amplification required development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative spatiotemporal analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive and are further strengthened in the revised version through additional analysis and adoption of new methods. This comprehensive analysis advances our understanding of MCC biology regarding the involvement of microtubules.

      Comments on revised version.

      The revised manuscript is substantially improved, and given the scope, it is appropriate that it primarily establishes a detailed spatiotemporal framework. That said, a few points would further strengthen clarity and impact. First, several observations naturally raise follow-up mechanistic questions, for example whether additional cytoskeletal systems such as actin contribute to steps like centriole apical migration. A slightly more detailed framing of these open questions would help guide future work. Second, some terminology introduced to label observed microtubule-based structures (for example "nest") may not be essential. Finally, while the authors have increased quantification, some analyses would benefit from super plot-style displays with replicate-level comparisons, particularly for intensity-based readouts.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully addressed the insightful comments provided by the reviewers which thoroughly increased our comprehension of the dynamics of centriole amplification. The manuscript has been revised accordingly and put in the context of the two papers we published since our last submission, showing that MCC differentiation is a genuine cell cycle variant. A point by point answer to all reviewer comments is provided below.

      Briefly:

      We have streamlined terminology and nomenclature in text and figures / better define experimental conditions with nocodazole

      We have tested the role of dyneins in the dynamics of centriole amplification

      We have done correlative light and electron microscopy on the early stages of centriole amplification

      We have analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors

      Collectively, this allowed us to make a clearer parallel with what occurs during centriole duplication and to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles.

      We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Public Reviews:

      Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      We have streamlined terminology and nomenclature, clarified the description of the complex events, and test the role of dyneins in centriole amplification. Microtubules density in MCC does not allow to extract information from imaging. In addition, we have done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Altogether, our new data allowed to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles. We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Reviewer #2 (Public Review):

      This important work will be of interest to centriole and cilia cell biologists. It describes in detail how microtubules control multiple aspects of centriole amplification in brain multiciliated cells. This study provides a greater time-resolved and molecular proteomic mapping of the different steps involved, with or without microtubule disruption. Boudjema et al. show that microtubules are important throughout the centriole amplification process, from the early stages, where the procentrioles emerge from a pericentriolar "nest", through the growth stage where microtubules maintain the perinuclear localisation, to the detachment stage, where microtubules assist in perinuclear disengagement and apical migration. The results are generally well supported by the evidence, but the manuscript would benefit significantly from some heavy editing to introduce more niche terms, standardize abbreviations in text, and labels on figures to help bring the readers, especially non-specialists, along with them - increasing the accessibility of their work.

      We thank the reviewer for his/her enthusiasm. We have streamlined terminology and nomenclature and clarified the description of the complex events to increase the accessibility of our work. We also replied points by points to his/her specific comments.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical-basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, and MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the CEN2-GFP; mRuby-DEUP1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described CEN2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of DEUP1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM, and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells (treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules in centriole amplification. Addressing the role of microtubules during different stages of centriole amplification required the development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive. This comprehensive analysis advances our understanding of MCC biology.

      Weaknesses:

      The role of microtubules and other molecular players during different stages of centriole amplification in brain MCCs can be further studied and strengthened using the tools developed in the manuscript. A more quantitative description of some of the analysis performed in the manuscript is required to strengthen the conclusions.

      We thank the reviewer for his/her enthusiasm. We have tested the role of dyneins in the dynamics of centriole amplification, done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Recommendations for the authors:

      As you will see, all reviewers felt that the analyses of the involvement of microtubules should be strengthened by including controls and additional experiments. Also, they agree that significant text editing would help to improve the manuscript's accessibility and readability.

      Specifically, they would suggest (1) streamline terminology and nomenclature in text and figures; (2) better define experimental conditions with nocodazole (concentrations used, effect on microtubules, effect on canonical centriole duplication); and (3), in the absence of other complementary genetic perturbation experiments, add a limitations paragraph in the discussion about conclusions drawn from nocodazole treatment alone.

      Reviewer #1 (Recommendations For The Authors):

      Main issues:

      (1) The authors use variable terminology to describe the same or similar events/structures. For example, in Figure 1 they refer to "centrosome stage" where they observe a pericentrin "cloud", which they later refer to as a "nest". In all other figures the first stage is not referred to as the "centrosome stage" but as the "cloud stage". Again, they also describe the "cloud" as a "nest" occasionally, but not always. In the cartoon, the nest is termed "centrosome cradle". The variable and inconsistent use of terms is confusing and the authors do not provide any explanation for the use of one vs. another.

      The text is now corrected. The centrosome stage corresponds to the stage preceding the beginning of centriole amplification in MCC progenitor. The pericentrosomal cloud of centriole and deuterosome elements forms later on, during the amplification A-stage. The formation of this cloud marks the beginning of A-stage, and persists up to G-stage where it dissolves. When we show that the cloud hosts the first stages of centriole biogenesis, we defined it as a “nest”. We do not use anymore the term craddle.

      (2) What prompted the authors to use the term "nest"? It gives the impression that they describe aspecific physical entity/structure (also depicted in this way in Figure 3P, with microtubules outside of this structure), but what is the evidence for this?

      The cloud is the spatial entity and the term “nest” is used to define a function of this transient compartment. We decided to keep the term “nest” as we now identified it with correlative light and electron microscopy, in addition to U-ExM, and show that the accumulation of centriole and deuterosome elements is accompanied by the formation of immature procentrioles, deprived of MT walls, as well as immature and empty deuterosomes. The scheme with MT outside the cloud/nest is misleading as we see MT organized by the mother centriole. We have now changed this.

      (3) The "nest" may simply be a dynamic accumulation of precursor particles around the centrosome, similar to what has been described for centriolar satellites. Rather than proposing a new entity, I suggest testing whether the "nest" particles may colocalize with PCM1 and thus may be related to centriolar satellites. Based on the data, the nest would simply be the centrosomal MTOC that organizes a radial microtubule array on which particles move around its center. In the absence of other evidence, I am not convinced that a new term is needed.

      We totally agree with the reviewer: the centrosome, as MTOC, concentrates centriolar and deuterosome components. This cloud is consistently dissolved when MT are depolymerized or dyneins inhibited. So, the physical entity is a “cloud”. We used the term “nest” to propose one function for this cloud which is to form deuterosomes and centrioles, before they move away for maturation. In fact, deuterosome and centriole formation are hindered when the cloud is dissolved. We have tried to edit the text all over the manuscript to make it clearer.

      (4) Role of MTs: are microtubules required or do they just facilitate some of the investigated events?

      The reason why the role of MT has not been tested yet during centriole amplification is probably because MT not only constitute the cell cytoskeleton on which molecular motors ride to transport cargos or distribute forces, they are also the core component of the structures we are studying. This is why we have tested a range of nocodazole concentrations and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (Fig. 4 Supplementary 1A-B). This may lead to an underestimation of the role of MT but we cannot study the role of MT on centriole amplification if centrioles cannot be formed.

      Does multi-ciliation in these models eventually occur normally under the concentrations and treatment conditions used here? This should be tested and discussed in the context of whether microtubules are indeed required and at what step of the entire process (amplification, migration, ciliogenesis) they may be critical.

      We did both chronic and acute treatments.

      Chronic treatments were done to test the overall efficiency of centriole amplification when MT (or dyneins) are perturbed. Chronic treatments were used to assess the role of MT (or dyneins) on the global efficiency of centriole and deuterosome formation (number of cells able to amplify, number/size/loading of deuterosomes, final number of centrioles (Fig. 4H-I, Fig. 4 Supplementary 2 B-D). In these chronic treatment, we focused on centriole amplification and not ciliation since it was the scope of this study. Also, we did not take ciliation as a readout of amplification because ciliation is relying on MT polymerization.

      Then, we also did acute treatments to test the role of MT (or dyneins) at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration; Fig. 4, 5, 7, 8 and associated supplementary figures). Since one stage is dependent on the precedent one, this enabled us to decipher the direct role of MT (or dyneins) on each single stage. We have now edited text, methods, legends and pictograms to be clear on whether acute or chronic treatment was done.

      (5) Can the authors include control (non-amplifying) progenitors in their analyses? It would be useful to know what the signal and distribution of each specific marker are before differentiation begins (before the cloud stage).

      Non amplifying progenitors are analyzed and constitute the so-called “centrosome stage”. We have now precised it and called it the “progenitor stage”.

      (6) Figure 2: Again, the terminology is confusing, since the authors describe that DEUP1 forms a "cloud" with centrin during the A stage.

      Corrections have been done as explained in point 1.

      (7) Description Figure 3: the authors introduce yet another term: "halo" A-stage. Is this the early A stage? Again, this is not explained and confusing. More systematic and consistent description is needed.

      Corrections have been done as explained in point 1. The term halos is used un the lab as it was the first term we used in our Nature paper in 2014 in reference to the halo described by Erich Nigg when they overexpressed Plk4. It was an error to use it in the manuscript.

      (8) Nocodazole treatments: the used concentrations are quite high.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely (Fig. 4 Supplementary 1A-B).

      (a) To avoid non-specific effects the authors should test what the minimal concentration is that completely depolymerizes microtubules in their cell model and perform analyses at this concentration.

      We have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A-B), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). In case it was not clear, we refer to this now several time and more clearly in the text and methods.

      (b) They should demonstrate depolymerization of microtubules by microtubule staining in the acute and chronic noc treatments and at the different noc concentrations used.

      This is, and was, in supplementary material (same, Fig. 4 supplementary 1A).

      (c) The authors should demonstrate that the used nocodazole concentrations do not impair normal centriole biogenesis during the cell cycle in these cells; if so, impaired assembly of centriole wall MTs may contribute to the observed effects in Figure 4.

      As mentioned in point 8b, we have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). The ability of the cells to form centrioles during chronic treatments were always assessed using immunostainings of SAS6 and/or CEN2-GFP signals (now exemplified in Fig. 4 Supplementary 1B). We also did EM analysis on cells treated with the highest doses of nocodazole (Nocodazole 10 uM for 24h) and this showed that centrioles can form with, what seems to be MT walls, in cells totally deprived of cytoplasmic MT fibers (Fig. 4 Supplementary 3-4). However, this does not show that all the cells can, because the number of cells that can be analyzed by EM are not sufficient to conclude. Also, one cannot assess whether MT walls are properly polymerized. However, the absence of MT walls should not change the results of the Figure 4, which are based on DEUP1, SAS6 or CEN2-GFP signals for deuterosomes and centrioles. Also MT depolymerization affects the formation of deuterosomes, which should not be altered by MT wall defects as it is not affected, even when centriole formation is blocked (LoMastro et al., 2024). Last but not least, we now show that blocking dyneins, as a comparable and even greater effect, on the formation of the cloud, deuterosomes and centrioles (Fig. 4C-I and Supplementary Fig. 4), which confirms that MTOC function, rather that MT wall formation, explain the centriole biogenesis alteration shown in Figure 4.

      (9) The authors repeatedly refer to the centriole-to-centrosome conversion of amplified centrioles and how this resembles centriole-to-centrosome conversion during the cell cycle. However, they incorrectly claim that this occurs at the G2/M transition. PLK1-dependent modification occurs at this stage, but conversion and PCM recruitment only occur after mitosis (see original work by the Tsou lab, which needs to be cited here).

      We agree with the reviewer. We have now added additional data to show clearly that centriole biogenesis, which requires two cell cycles to proceed in cycling cells, is accelerated during the MCC cell cycle variant where the elongation and maturation cycles are superimposed. This is now clearly shown in Fig. 3, 5, 9 and discussed.

      (10) Figure 6H-J: the authors claim that at low noc concentration, more D-stage cells showed incomplete disengagement than in controls, but the effect is shown only for the highest 10 µM concentration. Do any eof the phenotypes in Figure 6 also occur at the lowest noc concentration (assuming it depolymerizes MTs)? Again, it is crucial to demonstrate this, to exclude unspecific effects not linked to MT depolymerization.

      An error was made on the figure (but not in the legend). In Figure 6, chronic treatments are at 1 or 5 µM. Only acute treatments were done using 10 µM. In both cases, MT are not entirely depolymerized in these experiments (Fig. 4 supplementary 1A).

      (11) Disengagement, Figure 7: The authors describe that DEUP1 signal spreads all over the cytoplasm and becomes diffuse during this process, but one cannot see a diffusive signal throughout cells in the figures.

      We pushed the contrast to make it clearer but the deuterosomes are still bright at this stage and it is difficult to have both signal clear (now in Fig. 6B). We have also changed the example in video (now video 19) to show it more clearly with DEUP1 channel alone.

      (12) Figure 7: localization of disengaged centrioles at microtubule "nodes" is not clear from the images. There are many centrioles and random colocalization may be expected simply based on the high number. Higher resolution and/or magnification and quantification would be needed.

      We have edited and now say that centrioles “colocalize” with MT which, since centrioles nucleate MT, seems normal. We agree that it could be random, but given the density of MT, and the number of centrioles, it does not seem opportune to us to quantify. We can just say that we never see centrioles is regions that are deprived of MT.

      (13) The term "diffusive" to describe slow centriole movements in Figure 8 suggests that it is not motor or force-dependent, but there is no evidence for that. Movement based on opposing forces could produce a similar result, but would not be considered diffusive.

      We agree. We have changed “diffusive” by “diffusive-like”.

      (14) The manuscript would greatly benefit from the analysis of some candidate motor activities that may drive the movement and migrations of centrioles in this system. This would support the importance of the microtubule network for the specific steps in these processes, and better define its role beyond "being required". Dynein may be a candidate or minus end-directed kinesins. Since chemical inhibitors are available, these types of experiments would be straightforward.

      We formerly tested ciliobrevin but had hard time because of the small stability of the drug. Since our submission to eLife, we tested dynapyrazol and dynarestin and found dynapyrazol very efficient in dissolving the Golgi, a good readout of dynein inhibition. We sought to test the role of dyneins, using dynapyrazol, on (i) the formation of the pericentrosomal cloud in A-stage, (ii) the oscillation of DEUP1+ structures during A-stage, (iii) the number, size, loading of deuterosome, (iv) the final number of centrioles, (v) the migration to the nuclear membrane and (vi) the final apical migration of centrioles. The results are now inserted in main and associated Fig. 4, 5, 7, 8, 9.

      (15) Discussion:

      "the role microtubules" lacks "of"

      This is now edited.

      "This lack is..." Lack of what?

      This is now edited.

      "reflexive link" - meaning of "reflexive" is not clear in this context

      We have removed it.

      In my opinion, the study does not identify a nest composed of DEUP1, PCNT, and Centrin2; it only shows that these components accumulate as particles around the centrosome, which functions as MTOC. Consequently, it seems that the "nest" does not exist when MT is depolymerized. One could consider the center of the centrosomal MT array as a nest in this context, but there is no evidence of a specific new structure as suggested by the way the term is used in the manuscript.

      This is what we want to say: the center of the MT array become a nest in this context. We do not state that there is a specific new structure. We just say that MT and dynein dependent concentration of centriole and deuterosome components exists and that this region nests the birth of centrioles and deuterosomes. Also, this compartment is restricted in time and space, which justifies to use a specific term. The MTOC exists in the progenitor cell, while this compartment, marked by DEUP1, Centrin, PCNT accumulation, appears at the beginning of amplification and grows during A-stage to be dissolved at G-stage when all the deuterosomes and centrioles have moved away.

      What is the evidence that "DEUP1 is a centrosomal protein before building deuterosome structures"? It would be good to refer to the specific experiment. Does DEUP1 localize at centrioles also in the absence of microtubules? If not, I would not consider it a centrosomal protein.

      We have removed this statement to avoid misinterpretation.

      "This reminds the centriole-to-centrosome conversion..." the sentence is missing an "of"; also, again the authors confuse the order of events during the cell cycle, where centrosome conversion occurs after completion of mitosis, not at G2/M transition.

      We have removed this statement to avoid misinterpretation. Also, see Point 9.

      "microtubule dependent nuclear migration" should be rephrased; it sounds as if the nucleus migrates.

      This has been changed

      The following discussion of disengagement being linked to association with the nuclear envelope and resembling the process in cycling cells is misleading. In cycling cells movement of centrioles along the nuclear envelope occurs at G2/M and drives centrosome separation (separation of centriole pairs) in preparation for mitosis, not centriole disengagement.

      We are now clearer. We compare centriole-loaded deuterosome organization around the nuclear membrane to the migration of new centrosomes during early prophase (Fig. 5F-H, Fig. 5 Supplementary 2G-K).

      Regarding the possibility that forces by microtubules generated by the daughter centriole drive disengagement also in cycling cells, I would argue that this is unlikely since the daughter centriole can only nucleate microtubules after disengagement has occurred (and conversion to centrosome/PCM recruitment). Once this happens, it may physically separate the disengaged centrioles, which is a different type of activity. Indeed, originally the term "disengagement" was coined to specifically describe the loss of the perpendicular engagement of daughter centrioles with their mothers (Tsou and Stearns, Nature, 2006).

      We have removed this statement to avoid misinterpretation. The perpendicular engagement is difficult to assess on deuterosomes but we do see by live imaging, that attachment changes during D-stage, before centrioles detach clearly from deuterosomes.

      "high resolutive" should be "high resolution"

      Edit done.

      "splitted" should be "split"

      Edit done.

      "Consistently, when the mitotic oscillator is dis-inhibited and cells enter pseudo-mitotic events, centrioles show clear and rapid cell-cycle like clustering" This sentence is not understandable without further explanation; what does mitotic oscillator refer to? What are pseudo-mitotic events? What is cell cycle-like clustering?

      We have removed this statement.

      Minor:

      (1) Abstract: "Centriole number must be restricted to two..." Since cells are born with two centrioles and have 4 centrioles (2 pairs) when they enter mitosis, this sentence is inaccurate.

      The sentence has changed.

      (2) Abstract: "reflexive link"; I am not sure what the term "reflexive" refers to?

      We have removed this statement to avoid misinterpretation.

      (3) Figure 1C, D: it should be described better that the larger magnification panels represent overlays of many cells and what marker they show. This is not obvious since the smaller single-cell panels always show two different markers. Also, it would be more useful to show also single cells in the magnified view. The overlay does not allow us to see if a marker forms a cloud or a single dot, which is as important as the cell-to-cell variation in distribution.

      We have clarified this in the text and the legend. The cell-to-cell variation cannot be estimated with the overlay, but the projection from several cells (number precised) allows to see that the signal is confined in a restricted region. Or not. Which is what we wanted to analyze.

      Related to the above, the authors say that pericentrin forms a cloud at the top left in panel D, but there is only one confined centrosomal dot in the single-cell panel.

      The sentence has changed.

      (4) Results, Figure 2F; video 4: The authors claim connection and disconnection of DEUP1 aggregates with centrosomal centrioles; can the authors comment on the spatial resolution including in z in this movie to support this claim? Can they exclude that the structures are in proximity of each other rather than "connected"?

      This is a single z-section of 500nm. The resolution in xy is 128nm/pixel. Given the sizes of deuterosomes and a mature centriole, and given the fact that we observed this dynamics in several cells in live, we can state that the structures are connected. This is consistent with deuterosomes frequently observed “kissing” the daughter centriole by EM in the present manuscript (Fig. 2D, Fig. 2 supplementary 3 and 4 and Fig. 4 Supplementary 3-4). One has to look carefully at the daughter centriole (marked “dc”) and span in on the serial sections to see the connected deuterosome (marked by a star): this is at very early stage and therefore it is small. We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both structures are frequently sticked to each others on tens of nanometers.

      (5) The term "dynamics" as used in the manuscript should be plural.

      It has been used plural, except when for “dynamic microtubules” and “dynamic attachment to the nucleus”, which we think is ok? We have not found any other singular uses in our manuscript.

      (6) Figure 5: what does "YL1/2 procentriole intensity" refer to in panel F? This should be the intensity of microtubule asters.

      This has been modified.

      (7) Figure 6 - supplement 1B: contrary to the claim in the text, one cannot see tight colocalization with the nuclear pore marker. This seems to be a very small subset of particles and even in those cases colocalization is not tight. Also, what is the relevance of nuclear pore colocalization?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. What we want to say is that there is a tight connection with the nuclear envelope as shown by the localization of NPC on the same z-section as centrioles. This is why we present a single z, to show that centrioles and NPC are on the same z-plane of 500nm. NPC are stained to outline the nuclear membrane. This is also clearly visible for G-stage centrioles in the XY plane. We have now added an entire z-stack on video 18.

      Reviewer #2 (Recommendations For The Authors):

      To improve accessibility of their manuscript, we would suggest making the following edits:

      (1) Define 'specialist' or 'niche' terms each time you introduce them, such as 'pericentrosomal nest', or 'flower-like structures'.

      This has been clarified.

      (2) Have a think about abbreviations, again ones that work for people outside the project- this paper uses 'PC' for 'procentriole' but for many 'PC' is 'Parental centriole' or Figure 6J talks about 'D total' or 'D partial', leaves readers confused.

      This has been clarified.

      (3) Standardize your abbreviations throughout particularly for your treatments- sometimes Noco sometimes, NOCO, or your imaging experiments sometimes Cen-GFP, sometime CEN2-GFP (Figure 7A, D vs. Figure 6) or DEUP1- mRuby, DEUP1-mRuby3 or mRuby3-DEUP1?

      We now use Nocodazole or Noco in the text and the figure respectively, CEN2-GFP and mRubyDEUP1.

      (4) About 10% of the population, including several key figures in this field, are red-green color blind. Although 4 colour fluorescence is difficult to get right for everyone, choosing palettes (especially for two colour panels) is inclusive. More so, greyscale or inverted monochrome images make it easier for everyone to visualize changes in localization, size, and intensity. Red on black small foci is particularly difficult to discern. For example, Figure 3 - more individual channels in grayscale with arrows to mc, dc, and cilia would be helpful - difficult to distinguish stainings.

      We thank the reviewer for this comment and for this recommendation of being more inclusive. We have done the changes.

      To improve the conclusions drawn, we suggest some revisions below:

      (1) Since the paper really hangs on it, a clearer description of the rationale for when, how long and how much nocodazole treatment was done is needed. The logic currently is difficult to follow seemingly random jumps 10x concentration are used. Microtubules control many aspects of cell biology and could be impacted. For example, I particularly found Figures 6D and H difficult to follow i.e. the timing for 6H seems off.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely. We have of course tested a range of nocodazole concentrations at the beginning of the study and shown the extent of MT depolymerization under each treatment. We used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4 reviewer 1). The level of perturbation of MT and consequences on centriole formation at the different timings and doses were done for each experiment and are exemplified in Fig. 4 supplementary 1A-B. This figure was already present in the first version of the manuscript but we have now edited text, methods and pictograms to clarify this.

      (2) Perhaps an extension of this point- in general how interdependent are the processes? If there is a defect at the nest stage, how much are the later defects secondary to this, or do MTs genuinely play direct roles at all stages or are these knock-on effects? How do the authors rule this out? Defects in the nest, lead to smaller and more DEUP1+ foci, with defects in concentrating procentriole factors and centrin, which lead to... For example, Figure 4B looks like centrin is reduced upon noco treatment? Does noco treatment affect Cetn2GFP levels globally? Individual channels grayscale would help visualise this better.

      See also our answer to reviewer 1 point 8c.

      The stages are indeed interdependent. This is why we did both chronic and acute treatments. Chronic treatments were done to test the overall efficiency of centriole amplification when MT are perturbed. We typically used low dose of 1µM because nocodazole remains 48h in the culture medium. Acute treatments were done to test the role of MT at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration). Most of the acute treatments were done live and nocodazole was applied after the first time point of live monitoring. We used 10µM to have a rapid effect, and because nocodazole remains only several hours in the culture medium. This allowed to monitor the stage “n”, in cells where the stage “n-1” was completed without any drug which allowed to analyze a stage without having perturbed the precedent one.

      We now also test the consequences of dynein inhibition using both acute and chronic dynapyrazole treatments. We show that except for centriole migration, dynein inhibition phenocopies MT depolymerization (centriole number, perinuclear organization and disengagement as well as deuterosome number/loading/size).

      Nocodazole chronic treatments do affect intensity of CEN2-GFP at G-stage centrioles suggesting an altered A-to-G transition. In D-stage, CEN2-GFP signal seems normal. We now mention this in the text and in the Fig. 4 Supplementary 1B.

      (3) The authors nicely show the importance of MTs in the structure of the nest from which procentrioles and DEUP1 positive structures emerge. They suggest this nest may be what supports procentriole generation in the absence of DEUP1 and parental centrioles. Firstly how does this nest look in the absence of DEUP1 and/or parental centrioles (centrinone treatment)? This may be what they are trying to show in Figure 5 Supplement 1 but it currently is very difficult to digest what it is showing relative to controls and whether this is significant in the way it is plotted.

      The nest is conserved in the DEUP1KO with or without centrosomal centrioles, as shown by accumulation of Centrin and PCNT at the center of the self-organised MT network (Mercey et al., 2019). This is in fact what motivated our study on the role of MT in centriole amplification. We have edited the legend to precise the quantification done, which is not related to this question. In this quantification, we show that the increased propensity to accumulate PCNT by centriole-loaded deuterosomes between A and G-stage is maintained in the absence of deuterosomes, indicating that centrioles themselves accumulate/recruit PCNT.

      (4) Can you do CLEM on DEUP1-Ruby and these early foci at the cloud stage to see if they are visible at the ultrastructural level, relative to procentrioles, microtubules, and other electron-dense structures?

      We thank the reviewer for this question. We have done CLEM on the pericentrosomal cloud during very early steps of centriole amplification. This showed that DEUP1 early accumulation at the centrosome corresponds to a region rich in fibro granular aggregates, suggesting that DEUP1 may be translated here, through locally concentrated centriolar sattelites, known to be involved in local translation. Then, small deuterosomes and immature centrioles are formed, within this cloud of sattelites, confirming that the pericentrosomal cloud is a nest for centriole biogenesis (Fig. 2C-D + Fig. 2 Supplementary 2-6 for control and Fig. 4 Supplementary 3-4 for nocodazole treated cells). This also shows that immature deuterosomes are not necessarily round shaped, and can be deprived of centriole loading.

      (5) Check the scale bars- see Fig 4E. Check throughout.

      Done.

      (6) Figure 3 Supplement 1 and 2 don't match the legend and are likely reversed - which one is right?

      Done.

      (7) Technical issue - I couldn't play videos 6 or 16? Check these work.

      Done.

      (8) Nomenclature mammalian proteins- mouse or human- should be all caps DEUP1, PLK4, SAS6,etc. Watch your units- space between number and unit.

      This has been done.

      (9) Many of the graphs involve three biological replicates but why not plot the mean of each of the three experiments and do stats? The number of events measured may conflate the significance. Try using Superplots.

      Here is how we proceed: we count the number of occurrence of the phenotype we monitor, and the total number of cells. We apply a X<sup>2</sup> to test whether there is a significative difference between our replicates in each condition. If not, we pool the number of occurrence of the phenotype we monitor and the total number of cells for the 3 replicates, and for each condition. Finally we apply a X<sup>2</sup> between the different conditions. This is how we usually proceed to avoid comparing a mean of percentages. This is now explained in the methods.

      Minor points:

      (1) "DEUP1 is a centrosomal protein and assembles deuterosomes in the pericentrosomal region in brain MCC". I am not sure you have evidence that DEUP1 is a centrosomal protein. You don't seem to study the relationship between centrosomes and DEUP1? Rewrite this title and tone down this claim.

      This has been modified.

      (2) Why the crossbow micropattern (versus some other shape) - seems very specific but not discussed?

      We wanted a shape where centrosome is not localized at the center of mass of the nucleus. Among the corresponding patterns, the crossbow was the one where differentiating cells had less propensity to detach.

      (3) Figure 2 - are the foci of DEUP1 at the cloud stage smaller than at A stage? How do they grow? Measure the diameter at cloud stage, just after they leave the cloud and then once they move away from centrosomal cloud and each other. If so, and they do indeed grow in size from the cloud stage to the growth stage which I think your images suggest - do you envision this happening with the gradual addition of DEUP1 rather than fusion?

      Early deuterosomes are not easy to detect by light microscopy, because of accumulation of DEUP1 in the cloud. We did CLEM on the cloud of early A-stage cells to resolve the earliest deuterosomes which are often very small (see Fig. 2D, Fig. 2 Supplementary 2-6) suggesting that they grow, either by fusion, which we never observe in our movies at later A-stage, or by accretion of DEUP1. However, by light microscopy, we can detect very early but big deuterosomes, which we see splitting later on into smaller ones. So, we cannot conclude on the mechanism that regulate deuterosome size. This is now discussed in the discussion of the manuscript.

      You say in the discussion:

      "Consistently, we never observed fusion events of DEUP1 condensates in our time-lapse experiments. More importantly, we did FRAP experiments on endogenously tagged mRuby-DEUP1 in cells at the different stages of centriole amplification, and did not find significant recovery, supporting that centrosomal DEUP1+ foci and deuterosomes are not liquid-like structures (Figure 8 Supplementary 2)." How do you prove there is no fusion of deuterosomes?

      It is always difficult to prove the absence of something, we agree! But we did tens of movies with high temporal resolution and never observed fusion events. But, as we say in the previous question, the very early deuterosomes can be very small and we do not distinguish them from the DEUP1+ cloud by live imaging. So at this stage, we cannot say. But later on, during A- or G-stage and when deuterosomes are outside the cloud to be easily observed, we very often observe deuterosomes bumping into each others and stay in close contact for minutes, but then moving away. This, for us, supports the lack of fusion properties. But the question remains open. We now explain this in the manuscript and have added an example in video 28.

      If they are getting bigger as I think your imaging suggests from cloud to growth stage, then how is this happening?

      MT depolymerisation and dynein inhibition leads to the formation of very small deuterosomes. Dynein inhibition can even lead to a block in the formation of new deuterosomes suggesting that DEUP1 concentration is a crucial parameter for condensation into deuterosomes. Deuterosome growth may happen through oligomerization of DEUP1 molecules allowed by their dyne-independent concentration. Sorokin in 1968 proposed that a supersaturation of deuterosome components may lead to their solid crystallization into deuterosomes. Deuterosome size can also be regulated by a more complex molecular cascade, involving post-translational modifications of DEUP1 or PCM, such as phosphorylations driven by the cell cycle machinery. This would be consistent with the fact that deuterosomes are very big in the absence of CCNO, a cyclin required for entering the MCC cell cycle variant. This will need further investigations.

      I'm not sure FRAP actually proves fusion doesn't happen.

      Agreed, this is not what we wanted to say, we clarified. The FRAP experiment just suggests that it is not liquid-like.

      It is technically difficult to laser ablate individual or only subsets of deuterosomes...

      This is what was done but anyway, FRAP does not firmly show that deuterosome compartments are not liquid-like as we now precise.

      (4) How do you fix your cells for expansion as you have no preservation of cytoplasmic microtubules? You are saying that there is a "nest" of MTs but beta tubulin ONLY stains the cilia and centriole - why is this? Tyrosinated tubulin on regular confocal shows strong cytoplasmic staining. See Figure 3.

      Cytoplasmic microtubules do not preserve well through the expansion process. We did try a few different fixations and pre-extraction methods but they come at a trade-off to preserving centrioles. i.e. we could either preserve cytoplasmic tubes or centrioles but not both with the same processing method.

      (5) "PCNT puncta partially overlap with centrin (Figure 3 Supplementary 2C). At this stage, PLK4, the master regulatory kinase, and SAS6, one of the first centriolar components are either absent or present as small foci within the cloud, often on the wall of the parent centrioles (Figure 3B-C)." some arrows to highlight this would be useful - difficult to see?

      We have tried to make arrows on what is now Fig. 3 Supplementary 1 G, but there is to many CENTRIN colocalizing with PCNT. We have enhanced the contrast of the merge to make it more visible.

      (6) Figure 3I legend - what are the arrows pointing at? Yellow and white on inserts? ". Around the same time as tubulin, centrin is also recruited to procentrioles (Figure 3I). This stage is probably the stage that we previously documented as A"

      However you see centrin at DEUP1 foci in D, and you don't show any eg. SAS6 or PLK4 positive DEUP1+ structures lacking centrin specifically, centrin seems to be present on all the procentrioles in Figure 3I. Did I miss it where you show centrin negative procentrioles in the cloud?

      Fig. 3I (now Fig. Supplementary 1J), yellow arrows are pointing at centrioles with non-acetylated MT while white arrows point at acetylated MT. This is now indicated in the legend.

      Regarding CENTRIN, it is present as a diffuse staining around the centrosome since the very beginning of amplification (now in Fig. 3 Supplementary 1A with different contrasts), in addition to compose the parental centrioles. This staining can therefore overlap with DEUP1 staining when DEUP1 appears (Fig. 3 Supplementary 1B, E) but not necessarily. In live we observe that CENTRIN and DEUP1 foci can move independently at early stages (Fig. 2 Supplementary 1B, video 2). This is later on, as shown now in Fig. 3 Supplementary 1J (previously Fig. 3I), that procentrioles are all strongly positive for CENTRIN.

      A new paper (Laporte et al., Cell 2024) recently showed that the recruitment of CENTRIN on duplicating procentrioles first occurs at the distal end, visible by a small dot, and then appears gradually at the level of the inner scaffold when procentriole reach 160nm, the stage where POC5 appears, which corresponds to the A-to-G transition in our MCC progenitors (Al Jord et al., 2014). One can therefore consider that the same is happening in our cells, and that, with the CENTRIN cloud, we have difficulties to detect the distal CENTRIN dot. We have changed the text to add this reference and discuss CENTRIN apparition in MCC procentrioles.

      (7) " The DEUP1 asymmetry previously described at the centrosomal daughter centriole (Al Jord etal., 2014) becomes visible in some cells during the cloud stage (Figure 3B, N; Figure 3 Supplementary 2B) and in a majority of cells" difficult to see - maybe enlarge and single channel from Figure 3F-H in the supplemental Figure 3 to emphasise this?

      We have either changed the pictures or the contrast to be more representative with the quantifications. This is visible in Fig. 3A, D, E, G; Fig3. Supplementary 1E and now using correlative light and EM in Fig. 2 Supplementary 2, 3, 4 and Fig. 4 Supplementary 3-4. One has to look carefully at the daughter centriole (marked “dc”). We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both strutures are frequently sticked to each others on tens of nanometers.

      (8) Do you have videos of DEUP1 oscillations with nocodazole to show a lack of oscillations?

      We have now added videos of DEUP1 oscillations under nocodazole and dynapyrazole treatments.

      (9) "In addition, co-staining of centrioles and nuclear pore proteins show a tight colocalization(Figure 6 Supplementary 1B)." I see the colocalisation in panel 1 but less obvious with panel 2 maybe have some more zoomed in panels and some quantification of the colocalization? Is it more striking at the G stage than the D stage?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. There is too many centrioles and NPC, they cannot do otherwise than colocalize… What we want to say is that there is a tight connexion with the nuclear envelope. This is why we present a single z, to show that centrioles and NPC are on the same z-plane. This is also clearly visible for centrioles that are loaded on deuterosomes that are around the nuclear membrane in the XY plane. We also added a video to show an entire z-stack of this kind of staining.

      (10) "Indeed, SAS6 normally disappears from procentrioles when centrioles are docked, just beforeciliation (Al Jord et al., 2014). This suggests that centrioles were able to degrade SAS6, a process also dependent on APC/C (Strnad et al., 2007), but failed to disengage from deuterosomes." Figure 6 Supplement 1E-F - are you sure it wasn't that Sas6 wasn't loaded correctly at the earlier stage and so is reduced recruitment rather than premature disengagement of Sas6? If it is indeed premature disengagement of Sas-6 - what about CP110 - does the CP110 get loaded and is it still present in noco treated cells arrested in the D phase?

      We do not observe SAS6-negative procentrioles on deuterosomes at G-stage but only on deuterosomes in D-stage cells (cells with partly disengaged procentrioles). This is why we hypothesize that, because of the long duration of D-stage and knowing that SAS6 is finally degraded at the end of amplification (Al Jord et al., 2014), we are in the presence of cells where SAS6 has been degraded but where centrioles did not manage to disengage. This is now clarified in the text.

      (11) Can you track deuterostome splitting live? Maybe not enough spatial or time resolution?

      One has to monitor in 3D (multiple z because deuterosomes move a lot), 2 colors, high temporal resolution (dt=2-5’; to be able to track a single deuterosome), and long duration (deuterosomes are sometimes touching each other and then moving away, giving the impression that they split). This eventually leads to the bleaching of the mRuby fusion protein… We have put an example of what we think is a deuterosome splitting in Fig. 6E (former Fig. 7D). But we decided to finally monitor with low temporal resolution (dt=40’) to avoid photobleaching, and analyze numerous deuterosomes and cells to quantify the number and size of deuterosomes over time in single cells.

      (12) The MT nodes - can you segment the tyrosinated MTs and define nodes and then quantify theDEUP1 presence on them?

      Please see answer to reviewer 1 regarding this point.

      (13) Figure 8 supp 1 (E): Representative XY distribution of CEN2-GFP+ centrioles at the end of migration (Sas6 negative) in brain MCCs treated with DMSO, Nocodazole 1µM and 5µM (48h). Scale bar, 5µm Bit more detail on how you define fully migrated vs still migrating centrioles in z. You say you are using Sas-6 negativity to define fully migrated cells in the legend, yet you say noco treatment leads to premature sas-6 negativity, and yet the apical migration takes longer upon noco treatment?

      Nocodazole does not lead to premature SAS6 negativity but to a partial disengagement which lead to SAS6 negative “mature” centrioles being still connected to deuterosomes. We define complete migration when all the centrioles are on the apical side of the nucleus. We now clearly define what “apical” migration stands for in the main text and changed the pictograms in Fig. 8G to clarify this.

      (14) Figure 8H and video 18 - it isn't obviously clear to me that the noco-treated cells are "more erratic" or how you decide what counts as apically migrated successfully. How do you control for drift in z? Can you track individual centrioles as you did in untreated and define what is "erratic about their movement?

      Erratic means that the centrioles are moving away from each others, and back, in a non-predictable way, instead of migrating up and gathering. The drift in z of the whole cell is visible because there is always some centrioles, that are apically located at the beginning, that remains on the apical membrane, probably because they are already docked.

      We have indeed followed the centrioles individually in the nocodazole condition. However, in the control, the XYZ coordinates of one of the centrioles of the centrosome, which normally don’t move, are substracted to the coordinates of all the other centrioles as explained in the method section. This allows to have a subcellular reference, and to circumvent the movements of the cell, which are non-negligible at all at this timescale. In the nocodazole treated cells, the centrosomal centrioles share the erratic movements of the other centrioles and can migrate up and down, which exclude them as a reference. Since the nucleus is also moving a lot, we were left with no reference point.

      (15) Figure 8 supplement 1E can you quantify the final area of centriole patch in XY upon noco treatment?

      It was in main Fig. 8J and is now in Fig. 8 Supplementary 1F.

      (16) Figure 8J legend- MBB is never defined as an acronym.

      Thank you for pointing this.

      (17) Define what is the frequency and how is it calculated - Figure 8J.

      This is the MBB patch area in µm<sup>2</sup>

      Text edits:

      (1) "Altogether, these results suggest that, in this non-tissue-specific proxy of MCC progenitors, microtubules organize the onset of centriole amplification in the pericentrosomal region."

      Sentences have changed.

      (2) "Increasing the temporal resolution to 5-15s reveals that DEUP1+ foci observe an exhibit oscillatory dynamics to at the centrosome (Figure 2E, colored arrows, Video 3, 5/10 cells observed for 1-4min)."

      Sentences have changed.

      (3) "stage procentrioles were involved in this perinuclear migration and distribution. In fact, this dynamic is reminiscent of the centrosome migration that occurs during the G2-to-M progression in cycling cells in preparation for mitotic spindle organization. In cycling cells, this" Grammar - maybe change to "stage procentrioles were involved in this perinuclear migration and distribution. This is reminiscent of the centrosome migration that occurs during the G2-to-M".

      Sentences have changed.

      (4) "We then wondered whether these microtubule-dependent dynamics was were required for an efficient subsequent centriole disengagement during the following D-stage."

      Sentences have changed.

      (5) "Then, monitoring tens of disengagement movies, we identified a transient stage during which disengaging procentrioles redistribute isotropically in the 3 dimensions, along the nuclear membrane (Figure 6A, 4:30, Video 7) before losing its contact to migrate to the apical surface (Figure 6A, 6:30 to 14:00)."

      Sentences have changed.

      (6) Discussion: "Since pioneer electron microscopy studies on basal body production in quail oviduct MCC 35 years ago (Boisvieux-Ulrich et al., 1987, 1990; Boisvieux-Ulrich et al., 1989), this work is the first to assess the role of microtubules in the now finely described centriole amplification process. This"

      Sentences have changed.

      (7) "Using live imaging on brain MCC, we highlight the existence of a nest composed of DEUP1, PCNT and Centrin2, pre-assembled before the onset of centriole amplification onset."

      Sentences have changed.

      (8) "Recently, formation of DEUP1 pure condensates in solution as well as FRAP experiments after overexpression of DEUP1 in MCC progenitors suggested that deuterosomes where are not liquidlike structures (Yamamoto & Kitagawa, 2019). Consistently, we never observed fusion events of DEUP1."

      Sentences have changed.

      (9) "This reminds is reminiscent of the centriole-to-centrosome conversion occurring at the G2-M transition followed by the associated microtubule dependent nuclear migration of new centrosomes at mitosis onset (Agircan et al., 2014)."

      Sentences have changed.

      (10) "Following individual trajectories requires high resolutive resolution spatio-temporal live imaging while avoiding excessive light exposure which disturbs centriole migration (Boudjema et al., 2024)."

      Sentences have changed.

      (11) "Using high temporal resolution microscopy, we further identify that individual dynamics is are complex and can be splitted between divided into the baso-apical migration, where centrioles move in a processive and more..."

      Sentences have changed.

      Reviewer #3 (Recommendations For The Authors):

      (1) Growing MEF-MCCs on micropatterns has successfully mimicked the dynamics of centriole amplification in brain MCCs, allowing the authors to study the spatial origin of procentrioles. Since this is a powerful system, a more quantitative description of the system will be informative and beneficial for future studies. For example: What is the efficiency of this system? Do the cilia that form in MEF-MCCs motile?

      The system of MEF-MCCs has been described in a previous paper from the Kintner lab. It seems that growing the MEF-MCCs on micropatterns did not ameliorate the ciliation which is partial, probably due to the absence of an apico-basal polarity.

      (2) Figure 2: The analogy drawn by the authors between DEUP1 oscillatory dynamics and centriolar satellites is intriguing. In early amplifying cells within the cloud, do these DEUP1 structures co-localize with the satellite marker PCM1?

      We have added immuno stainings of PCM1 in mRuby-DEUP1 / CEN2-GFP cells in Fig. Supplementary 2E. Within the centrosomal cloud, DEUP1 colocalizes with PCM1. Interestingly, this PCM1 concentration at the centrosome is dependent, at least in part, on dyneins. Then, PCM1 can localize around the deuterosomes, but it is never colocalized with deuterosomes (not shown). This is also showed by immuno-EM in Zhao et al., 2019. Although it was shown that PCM1 is a proximity interactor of DEUP1 (called ccdc67 at that time) by Firat-Karalar et al., 2014., absence of PCM1 staining on deuterosomes does not favor the hypothesis of PCM1 and DEUP1 being part of the same entities. One could hypothesizes that DEUP1 is transcribed locally within the satellites, explaining the colocalization of the 2 proteins and the + BioID results, and then form PCM1negative deuterosomes.

      (3) The authors propose a physical link between deuterosomes and centrosomes based on their oscillatory behavior. How are the oscillatory dynamics of DEUP1 affected by nocodazole treatment or inhibition of microtubule motors (i.e ciliobrevin treatment)?

      These oscillations are inhibited by nocodazole (Fig. 4D). They are also inhibited by dynapyrazole (Fig. 4D). We never succeeded in having a nice disruption of the Golgi apparatus with ciliobrevin and therefore we did not used it.

      (4) In addition to nocodazole treatment, it would be important to determine the consequences of microtubule stabilization by taxol and inhibition of microtubule motors during critical stages of centriole amplification where microtubules are reported to play a role for the first time in this manuscript. Another interesting area of investigation will be to study the extent to which microtubule PTMs contribute to these processes.

      We now blocks dyneins during the different stages of amplification. The results are in main and associated Fig. 4, 5, 7, 8. The role of microtubule PTM, is not in the scope of this manuscript.

      (5) Describing microtubule dynamics along with Centrin/DEUP1 dynamics will be informative in assessing whether these structures associate and/or move along microtubules? Have the authors performed their imaging experiments with SIR tubulin?

      Yes, we have tried hard! But we have encountered different obstacles:

      3-color video microscopy is phototoxic,

      siRTubulin is bleaching very rapidly

      The density of microtubules in MCC makes the observation hardly informative

      (6) Figure 5: The role of PLK1 in centriole-centrosome conversion and generation of multiple MTOCs can be tested with a PLK1 inhibitor for further confirmation.

      We have also tried but inhibiting Plk1 blocks the A-to-G and G-to-D transitions so it was not possible to uncouple the role of Plk1 in stage transitions versus centriole maturation.

      (7) Figure 6: The tight co-localization of nuclear pore proteins with centrioles poses questions about the role of nuclear pore proteins or other nuclear proteins that are associated with centrioles during centriole disengagement and migration. Considering the existing literature on centrosome-nucleus attachments, can there be a way to test this question within the scope of this manuscript?

      We have tried to deplete Nup133 but it’s killing the cells. Our additional experiments now show that the nuclear migration of centrioles during G-stage is dynein dependent, reinforcing the parallel with centrosome migration in prophase. We also added results from our scRNA sequencing (Fig. 5 Supplementary 1) showing that some key players of centriole migration to the nuclear membrane are conserved in the MCC cell cycle variant, and expressed with a comparable dynamics as to the canonical cell cycle.

      (8) Figure 8: Manually tracking a subset of migrating centrioles to define their dynamics during centriole migration and docking provides valuable analysis for determining the molecular mechanism of these processes. In addition to microtubules, does actin contribute to this process? Since centrioles eventually migrate to the apical side in nocodazole-treated cells, there should be other molecular players involved in this process.

      We did block actin polymerization but we found that the different stages were affected and that it would be better to dedicate a whole manuscript on the role of actin during each stage of amplification. We discuss the migration mechanism, and the putative role of actin, in the discussion.

      (9) The legends for Supplementary Figures 1 and 2 in Figure 3 are mixed and need correction.

      Figures have been remodelled.

      (10) In Figure 3P, the term "PLK4+" is labeled in bright green, which is not clearly visible. It maybe beneficial to change the color of this label for better visibility.

      We have tried to correct this.

      (11) Figure 6F quantifies "% tethered flowers" on the nuclear membrane. When quantifying, is the3D localization of DEUP1 flowers in both DMSO- and Noc-treated cells considered? A flower may appear to be on the nucleus in 2D, but it could be detached from the membrane in a 3D view.

      The quantifications are done in 3D. However, flowers that are below or above the nucleus are not quantified since the space is confined and the resolution in z to small to see whether they are connected or not. This is now precised in the legend.

      Before the editors proceed with an updated assessment, they've requested that we pass on some of the comments that have arisen as part of the evaluation of your revised manuscript. They feel that these concerns should be addressed before we proceed with issuing a formal assessment and publishing the revised Reviewed Preprint:

      We thank the reviewers and the editors for the corrections and insighfull comments. We apologize for our delayed answer and hope our corrections in the main text and some of the figures will give them satisfaction.

      The revised manuscript is greatly improved with nice new data regarding the role of microtubules. It also has changed quite a bit including the title. The new focus is on the cell and centriole cycle variants in MCC. While this helped to focus the study, there remains an important issue related to the interpretation of the data and the proposed 2-in-1 cycle model. Before providing the final updated assessment, we ask you to address the following points (which were raised already in the first round of review): The manuscript still contains statements that are not aligned with published work and the current view in the field regarding the timing of events during canonical centriole biogenesis. These timings are in conflict with your model that 2 centriole cycles are "superposed" in the MCC cell cycle variant, as currently presented. An alternative straightforward interpretation would be that multiciliogenesis uses an accelerated centriole duplication cycle where key steps occur concomitantly or in short succession instead of being separated by mitotic divisions as in the canonical cycle.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we think we have proposed. When correcting our confusions as regard to centriole-to-centrosome conversion (as explained below) and putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (corrected Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a superposition of events that; although driven by the same molecular machinery, are normally occuring in two consecutive cell cycle. We explain ourself briefly in two paragraphs, before answering point by point to the questions of the reviewers.

      As regard to centriole-to-centrosome conversion:

      We thank the reviewer for pointing out that we used “MTOC conversion” for what is normally called “centrosome maturation”. We have removed the term “centriole-to-centrosome conversion” during the first round of revision but we now realize that “MTOC conversion” leads to the same misinterpretation as regard to the literature on centriole duplication.

      The reviewer asks us to refer to the work of the Tsou lab (Wang 2011, reference now added in the manuscript) showing that daughter centrioles are “modified” (e.g. recruit PCM, become competent for MT nucleation and duplication) during late M/early G1. This “centriole-to-centrosome conversion” can’t occur for our procentrioles at this stage since they are not even born during the mitosis that precedes MCC differentiation. Also, in our cells, such modification does not include the capacity to become competent for duplication since we know that procentrioles become basal bodies without making any round of duplication (Al Jord et al., 2014).

      Also, we have not done the experiments to tackle the question on when our centriole become “modified-like”. What we can say is that during A-stage, they become progressively positive for PCM (Fig. 5 Supplementary 2) and a weak signal shows that some MT are seen emerging from them (Fig. 5 and Fig. 5 Supplementary 2, and see point by point answer).

      What we do see is that, at the A-to-G transition, they increase their PCM recruitment, show clear and strong MTOC ability (sometimes as strong as the centrosomal centrioles), and that this is associated with migration and separation of centrosome/deuterosomes around the nuclear membrane (Fig. 5). We therefore connect this to what occurs at the G2/M transition which is an increased recruitment of PCM protein, an increased ability to nucleate MT, associated with centrosome migration and separation at the nuclear membrane. Since this process in the canonical cell cycle is called “centrosome maturation”, we therefore should refer to this term in our study. However, centrioles in the MCC variants are not organized in centrosomes, so we now compare what we see to the “centrosome maturation” of the canonical cell cycle with an associated reference (Joukov et al., 2018), but name it “centriole maturation”.

      We have modified the text (track changes visibles) and the schemes (Fig. 5, Fig. 5 Supplementary 1 and 2, Fig. 9, Fig. 9 Supplementary S1; new versions uploaded) accordingly.

      As regard to 1.5 or 2 cell cycles

      Except for the “MTOC conversion” that we have now changed, as explained above, we think our work does suggest (depicted on Fig. 9) what the reviewer states for centriole duplication: “In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs. Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium)”.

      We feel that going from early S to a G1 phase, after 2 mitosis, is what one can call “2 cell cycles”. One of the paper that inspired us a lot when studying how the cell cycle machinery can drive centriole amplification in MCC is a paper from Jadranka Loncarek team (Kong et al., 2014) where they also state that “nascent centrioles gradually mature through 2 cell cycles”. Very interestingly, in this study they show that when they enhance Plk1 activation, they could erase centriole age and new procentrioles are able to recruit PCM and appendages within only 1 cell cycle, without mitotic progression, like what we see in MCC. We have added the reference in our discussion.

      Point by point answer

      (1) Original work on canonical centriole disengagement and centriole-to-centrosome conversion should be cited (e.g. PMID: 16862117, PMID: 21576395)

      As explained earlier, we used the wrong term since the begining. We do not speak about the centriole-to-centrosome (nor MTOC) conversion since we do not test when centriole modification (Wang et al., 2011) occurs in the MCC cell cycle variant. We know that PCNT is present on the procentrioles during A-stage (as shown in Fig. 5 Supplementary 2B), but we do not know when it is recruited (UExM did not work properly with this antibody). We quantify a weak MT staining in regrowth experiment during A-stage and see that procentrioles can be connected to MT in both brain MCC and MEFs (as shown in Fig. 5D, E for brain MCC and Fig. 5 Supplementary 2F for MEFs) , but we do not know when during A-stage they become competent for nucleation. We therefore did not speak about this process that we do not document. What we clearly document/quantify is the enhanced MT nucleation capacities at the A-to-G transition, concomitent with the nuclear migration (easily defined with Cen2-GFP or GT335 stainings) and that we compare to centrosome maturation occuring at the canonical G2/M transition.

      (2) The authors state in several places that canonical centriole formation and maturation takes two iterations of the canonical cell cycle. This is imprecise. Based on the above work and work by others, the broadly accepted view is that it takes 1.5 cell cycles. This difference matters for the final proposed model (see below). Reviewed e.g. here: PMID: 20869612; PMID: 30601682

      Our answer is in the preamble.

      (3) "Centriole maturation cycle superposes with centriole elongation cycle in the MCC cell cycle variant": Your description of the canonical cycle differs from the current view in the field. In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). All this occurs in 0.5 cycles. Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs (total of 1.5 cell cycles). Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium).

      (4) Fig 5A, B and Fig. 9

      (a) Are 2 separate figures needed for the model? They seem redundant.

      We find it easier not to wait Fig. 9 to have the first part depicted.

      (b) The model shows loss of SAS6 throughout G1, but this already occurs during M/early G1

      Thanks. It was already ok in Fig. 9, we have modified for Fig. 5.

      The model shows "MTOC capacity/conversion" during S phase, but this occurs during early G1

      Thanks a lot, as explained earlier, we used the term MTOC conversion occurring in G1 for what is normally called centrosome maturation occurring in G2/M, as explained earlier. We do not speak anymore of MTOC conversion since we have not tackled this question (explained above). We have therefore removed MTOC conversion in the texts and the schemes and replaced it by “centrosome maturation” for the duplication cycle, and by “enhanced MT nucleation capacity” for the MCC cycle. To be clearer and schematize that procentrioles are competent for MT nucleation before G2/M or A/G transitions, we have added some MT nucleated from G1 procentrioles during the canonical cycle, and from late A-stage procentrioles during the MCC cycle.

      The model shows disengagement only in the second M phase, but this occurs already at the first M phase, directly following centriole biogenesis, right before centosome conversion.

      This is a big edition error in both Fig. 5 and 9. Of course the daughter centriole disengage during the first M-phase. This has been changed. Thanks a lot for spotting it. This, however does not contradict the hypothesis of superposition.

      We also added the acquisition of distal appendage which was written in Fig. 5 but not in Fig.

      9 for duplication during the second M-phase.

      When the correct timings are incorporated in the figure, the proposed superposition of two cycles is not an accurate description of the events. Instead, your data seem consistent with a model where MCC incorporates all steps in one cell cycle variant that lacks mitoses, so that disengagement and MTOC conversion occur together with centriole elongation, followed immediately by acquisition of DAs and SDAs.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we tried to propose. When putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a super opposition of events that; although driven by the same molecular machinery, are normally occurring in two consecutive cell cycle. This is notably consistent with the findings of Kong et al., 2014 cited previously.

      (5) While all reviewers felt that there was no need to introduce the new term "nest", they leave it to the authors to keep it. However, the authors may want to consider that the term is still not introduced and explained properly, which may confuse readers. For example, while this section reads like an introduction to the term: "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC", the term is already used two times before without explanation. The first mentioning is at the beginning of the results section and is followed by citations, which gives the impression that these studies describe the nest, which is not the case.

      The first mention of “nest” is in the end of introduction resuming the findings of the paper where the term is in the following context: “we found that centriole amplification emerges in a pericentrosomal “nest” concentrating core centriole/deuterosome elements”. We looked at nest definition in the Collins Dictionnary : “a structure or other place where creatures, esp. birds, give birth or leave their eggs to develop”, we felt this was clear. We added quotation marks around the term nest.

      Then, the result section opens with this sentence: “The origin of amplified centrioles in MCC remains controversial. Some live imaging experiments and electron microscopy suggest that the centrosome could constitute a nest for centriole and deuterosome biogenesis (Al Jord et al., 2014; Kalnins et al., 1972; Mori et al., 2017), but others have proposed that procentriole-loaded deuterosomes emerge independently from the centrosome location, all over the cytoplasm (Nanjundappa et al., 2019; Sorokin, 1968; Zhao et al., 2013, 2019).”. Here, the term nest is again used as a place of birth for centrioles and deuterosomes which is what is actually proposed in these papers. First, Kalnins el al., in 1969 (we made an error on the reference date, this has been changed), resume in their abstract “This observation suggests that all of the clusters may form initially in close association with the diplosomal centrioles”. Then, not to mention Al Jord 2014 which comes from our lab, the title of Mori et al. is “Cytoplasmic E2f4 forms organizing centres for initiation of centriole amplification during multiciliogenesis”, and in the paper, they show that E2F4 accumulates at the centrosome. This is now also proposed by collaborators for MCIDAS (Lu et al., 2025). We feel that these references, which are often omitted, are appropriated at this location.

      Then we continue with: “To test whether microtubules drive the organization of a centrosomal nest from which procentrioles emerge”, which keeps the notion of the place of birth.

      Then the title "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC" arrives. In this section we first speak about a pericentriosomal cloud on which we zoom in using CLEM, to then conclude at the end of the section “Altogether live imaging mRuby-DEUP1/CEN2-GFP during early A-stage suggests that core deuterosome and centriole components are concentrated in a primordial cloud around the centrosome, which constitutes a nest where centrioles and deuterosomes concomitantly form before they move away from the centrosomal region (Fig. 2F)”.

      Finally, we begin the discussion section regarding the nest by: “We named this transitory compartment a “nest” since deuterosomes and procentrioles emerge specifically in this region and grow while moving away from it.”

      During the first revision, we tried to make it clearer. If this is still not the case after and the reviewer has another proposition of definitions/phrasing, we will be glad to consider it.

      As replied to the other reviewer, the term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the transitory region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      The following comments from Reviewer #3 may also provide further context regarding the editors' remaining concerns:

      The authors have done an excellent job addressing the points I raised overall, and the revision is substantially improved in focus and clarity. That said, some concerns raised by other reviewers, particularly regarding terminology and statistical analysis, could have been addressed more fully. One issue remains insufficiently resolved. Several quantitative analyses (for example Fig. 5C and 5E) still appear to rely on pooled single-event measurements collected across three independent experiments. This approach can overstate statistical significance. The authors indicate in their rebuttal that they use chi-square tests to compare proportions and to justify pooling across replicates. However, I am not convinced this addresses the issue for the intensity-based and single event distributions shown in the panels specified above. I recommend that these key analyses be represented with biological replicates shown explicitly (superplot-style, with replicates distinguished).

      Our reply was for the comparison of proportions and not the intensity-based and single event distributions shown in the panels Fig. 5C and Fig. 5E. We have now changed our plots to represent biological replicates explicitly (superplot-style, with replicates distinguished). As for the statistical analysis: we evaluated differences in marker intensity between A-stage and G-stage samples using a linear regression model, with stages as the main effect and replicate as a fixed covariate, to account for batch variation. Statistical significance was assessed using Type II ANOVA.

      Separately, I continue to feel that some newly introduced terminology (for example, the "nest") may not be necessary at this stage. It may be sufficient to describe these structures and focus on their spatiotemporal behavior, composition, and measurable features, rather than assigning new names. Having read the authors' response, I understand that they would like to retain this terminology, which is acceptable; however, it may not be readily adopted by the field.

      The term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      Minor correction (remove "in MCCs" part from the following sentence):

      In MCC, PCM1 depletion alters deuterosome formation and centriole production in brain and airway MCC (Hall et al., 2023; Zhao et al., 2021).

      Done

    1. eLife Assessment

      This manuscript provides valuable insight into how genome organization changes as cells progress through the cell cycle after mitotic exit, identifying two sharp genome remodeling events at G1-S and to a lesser extent, at S-G2 transitions. The conclusions are supported by solid, rigorous data, including sequencing and orthogonal imaging data. The use of sorted unsynchronized cells rather than cells treated with drugs is a particular strength.

    2. Reviewer #1 (Public review):

      This work convincingly shows that, rather than gradually "evolving" throughout interphase, global chromatin architecture undergoes unexpectedly sharp remodeling at G1-S (and to a lesser extent, S-G2) transitions. By applying "standard" Hi-C analyses on carefully sorted cells, the authors provide an excellent temporal view of how global chromatin architecture is changed throughout the cell cycle. They show a surprisingly abrupt increase in compartmentation strength (particularly interactions between the "active" A compartments) at G1-S transition, which is slightly weakened at S-G2 transition. Follow-up experiments show convincingly that the compartment "maturation" does not require the DNA synthesis accompanying S phase per se, but the authors have not identified the responsible factors (work for future publications). The possible biological ramifications of these architectural changes (setting up potential replication "factories", and/or facilitating transcription-replication conflict resolution, both more pertinent for the active A compartments, which are most affected) have been well discussed in the article, but still remain speculative at this stage.

      My major criticism of this article is aimed more at the state of the field in general, rather than this specific article, but it should be discussed to give a more balanced view: what actually is a chromatin compartment? Chromosomal tracing and live tracking experiments have shown that the majority of "structures" identified from Hi-C experiments are statistical phenomena, with even "strong" interactions only being infrequent and transient. A-B compartments are "built up" from multiple very low-frequency "interactions", so ascribing causal effects for genome functions is even tougher. As a result, I have very little confidence in the results of the authors' polymer simulations and their inferred "peninsula" A compartment structures without any other supporting experimental data.

      Comments on revised version.

      The authors have included orthogonal DNA FISH evidence to support their claims which greatly strengthens the manuscript. Their further precisions within the discussion have answered all of my previous concerns with the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript by Choubani et al presents a technically strong analysis of A/B compartment dynamics across interphase using cell-cycle-resolved Hi-C. By combining the elegant Fucci-based staging system with in situ Hi-C, the authors achieve unusually fine temporal resolution across G1, S, and G2, particularly within the short G1 phase of mESCs. The central finding that A/B compartment strength increases abruptly at the G1/S transition, stabilizes during S phase, and subsequently weakens toward G2 challenges the prevailing view that compartmentalization strengthens monotonically throughout interphase. The authors further propose that this "compartment maturation" is triggered by S-phase entry but occurs independently of active DNA synthesis, and that it involves a consolidation and large-scale reorganization of A-compartment domains.

      Strengths:

      Overall, this is a thoughtfully executed study that will be of broad interest to the 3D genome community. The data are of high quality, and the analyses are extensive, albeit not completely novel. In particular, previous work (Nagano et al 2017 and Zhang et al 2019) has shown that compartments are re-established after mitosis and strengthened during early interphase, and single-cell Hi-C studies have reported changes in compartment association across S phase. In particular, Nagano et al show that DNA replication correlates with a build-up of compartments, similar to what is presented here, with the authors' conclusion that compartment strength peaks in early S. The idea that it weakens toward G2, rather than continuing to strengthen, appears to be novel and differs from the prevailing framing in the literature.

      Comments on revised version.

      The authors have responded constructively to my major conceptual concerns. The distinction between DNA synthesis and replication initiation has been clarified appropriately. The additional insulation analysis substantially strengthens the argument that compartment maturation is not simply a consequence of changing loop extrusion dynamics, although I would encourage slightly more cautious wording regarding "independence" from cohesin-mediated extrusion. The peninsula model is now framed appropriately as a heuristic interpretation and supported by orthogonal imaging data. Finally, the discussion of conservation across cell types has been appropriately tempered. Overall, I believe the manuscript has been significantly improved.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work convincingly shows that, rather than gradually "evolving" throughout interphase, global chromatin architecture undergoes unexpectedly sharp remodeling at G1-S (and to a lesser extent, S-G2) transitions. By applying "standard" Hi-C analyses on carefully sorted cells, the authors provide an excellent temporal view of how global chromatin architecture is changed throughout the cell cycle. They show a surprisingly abrupt increase in compartmentation strength (particularly interactions between the "active" A compartments) at G1-S transition, which is slightly weakened at S-G2 transition. Follow-up experiments show convincingly that the compartment "maturation" does not require the DNA synthesis accompanying S phase per se, but the authors have not identified the responsible factors (work for future publications). The possible biological ramifications of these architectural changes (setting up potential replication "factories", and/or facilitating transcription-replication conflict resolution, both more pertinent for the active A compartments, which are most affected) have been well discussed in the article, but still remain speculative at this stage.

      We thank Reviewer #1 for their positive and constructive assessment of our work, and we agree that the questions of responsible factors and biological ramifications are important directions for future studies.

      My major criticism of this article is aimed more at the state of the field in general, rather than this specific article, but it should be discussed to give a more balanced view: what actually is a chromatin compartment? Chromosomal tracing and live tracking experiments have shown that the majority of "structures" identified from Hi-C experiments are statistical phenomena, with even "strong" interactions only being infrequent and transient. A-B compartments are "built up" from multiple very low-frequency "interactions", so ascribing causal effects for genome functions is even tougher. As a result, I have very little confidence in the results of the authors' polymer simulations and their inferred "peninsula" A compartment structures without any other supporting experimental data.

      We thank the reviewer for raising this important conceptual point. This issue extends beyond the scope of the present study but reflects an important ongoing discussion in the 3D genome field regarding the biological interpretation of chromatin compartments.

      We agree that Hi-C interactions should not be interpreted as stable pairwise contacts present in every cell. A growing body of evidence from chromatin tracing and live-cell imaging studies has demonstrated that many chromatin interactions identified by Hi-C are probabilistic and dynamic, with substantial cell-to-cell variability. Relatively speaking, however, A/B compartment organization represents a robust population-level property of genome organization that is highly reproducible across biological replicates and closely correlates with multiple independent genomic features. In particular, replication timing (RT) correlates very well with A/B compartment organization, with early and late RT domains corresponding to A and B compartment domains, respectively.

      Furthermore, single-cell DNA replication sequencing (scRepli-seq) analyses have revealed remarkably low cell-to-cell variability in RT, suggesting that RT profiles and A/B compartment organization reflect biologically meaningful and relatively stable features of nuclear architecture rather than purely statistical artifacts. Thus, while individual chromatin contacts may be transient and probabilistic, the megabase-scale compartment organization inferred from them appears sufficiently reproducible to support reproducible RT programs and other genome functions. Additional support comes from decades of work on DNA replication demonstrating that spatiotemporal replication patterns, visualized as replication foci following short EdU pulses, are remarkably reproducible between individual cells throughout S-phase progression. These patterns reveal clear spatial segregation between early-replicating A-compartment regions and late-replicating B-compartment regions even at the single-cell level.

      To directly address the reviewer’s concern that A/B compartment organization might represent only an ensemble-level statistical phenomenon without biological relevance at the single-cell level, we performed L1/B1-EdU DNA FISH on asynchronous mESCs and MC12 embryonic carcinoma cells. L1 elements are enriched in B compartment domains, while B1 elements are enriched in A compartment domains, allowing visualization of compartment segregation in individual nuclei across the cell cycle. This single-cell analysis confirmed our Hi-C findings: compartment segregation increased from G1 to early S, remained elevated throughout S phase with reduced cell-to-cell variability, and then weakened in G2. Thus, compartment segregation is detectable in single cells, and the temporal dynamics of compartment maturation identified by population Hi-C were independently recapitulated at single-cell resolution. We have added a new Results section describing these findings titled “Stepwise A/B compartment reorganization during interphase is conserved at single-cell resolution”, including new Figure panels 2D–H and Figure S5.

      Regarding the polymer simulations, we agree that these models should be interpreted with caution. We do not view them as direct representations of individual nuclei, but rather as heuristic models that help visualize structural trends present in the Hi-C data. To make this point explicit, we have added the following statement to the revised manuscript: “We note that these models are derived from population-averaged Hi-C data and should therefore be interpreted as a heuristic framework for understanding A/B compartment dynamics, rather than as definitive representations of individual nuclei.”

      That said, we did try to provide orthogonal experimental support for the "A peninsula" model by performing DNA FISH. In brief, we measured distances between probe pairs spanning two A domains on chromosomes 2 and 15 across different cell-cycle stages. We observed significant increases in inter-probe distances from G1 to early/mid S, with the most pronounced changes involving the central probes (i.e., probes located near the domain center), consistent with physical extension of the A domain during S phase. While these data do not prove the exact geometry depicted by the model, these findings provide independent experimental support for the peninsula model as a simplified but biologically grounded interpretation of the Hi-C data. These results are described in the Results section titled “A-compartment consolidation during S-phase involves enhanced long-range contacts and structural reorganization” and are presented in new Figure panels 5D–F and Figure S12.

      We thank the reviewer again for raising this important conceptual issue, which prompted us to better clarify both the biological interpretation and the limitations of our analyses.

      Specific minor points:

      (1) A better explanation for how Figure 1E was generated is required, because this figure could be very misleading. Figure 1F and all other cis-decay plots (and the Hi-C maps themselves) show that the strongest interactions are always at smaller genomic separations, so why should there be more "heat" at the megabase ranges in Figure 1E?

      We appreciate the reviewer's observation. The apparent discrepancy is simply due to the fact that the decay plot (Fig. 1E in the original submission, now Fig. S2C) does not include the shortest-range interactions. The lowest distance plotted is 25 kb, following the method originally described in Nagano et al. (Nature, 2017), which we used as a reference. The shortest-range interactions (below 25 kb) are indeed the most enriched, as seen on the diagonal of the Hi-C maps (Fig. 2A) and in the standard cis-decay plot (Fig. 1F in the original submission, now Fig. S2F). With the 25 kb cutoff in place, the "heat" observed at megabase distances (specifically 12–50 Mb) in early/mid G1 corresponds to the dark, non‑specific band around the diagonal visible in the Hi-C maps at the same time points. This is also reflected in the cis-decay plot (Fig. S2F), where distances in that range appear above the expected curve (a "bump" rather than a linear decay).

      To avoid confusion, we have updated the figure legend accordingly (Fig. S2C): “(C) Contact decay profiles for all cell cycle phases, plotted from 25 kb to 50 Mb, illustrating a continuum of cis-interactions and a progressive shift from long-range (> 12 Mb) to short-range (< 1 Mb) interactions during the G1-to-S phase transition.”

      We hope this explanation clarifies the figure.

      (2) An ultra-high-resolution Hi-C study (Harris et al., Nat Commun, 2023) identified very small A and B compartments, including distinctions between gene promoters and gene bodies, raising further questions as to what the nature of a compartment really is beyond a statistical phenomenon. It is unreasonable to expect the authors to generate maps as deep as this prior study, but how much do their conclusions change according to the resolution of their compartment calling? The authors should include a balanced discussion on the "meaning" of A/B compartments.

      We thank the reviewer for highlighting recent ultra-high-resolution work, such as Harris et al. (Nat Commun, 2023), which reveals compartment-like features at much finer genomic scales. We agree that these findings raise important questions regarding the scale-dependence and interpretation of A/B compartmentalization.

      In our study, we specifically focus on coarse-grained compartment organization, analyzed across multiple resolutions (from ~1 Mb to sub‑megabase scales). Importantly, the key conclusions, including the abrupt strengthening of compartmentalization at the G1/S transition, are robust across these resolutions.

      We also note that fine-scale compartment-like features likely operate under different rules than larger-scale compartments. Recent evidence suggests that these "micro‑compartments" are more dynamic and transient (Harris et al., Nat Commun, 2023; Goel et al., Nat Struct Mol Biol, 2025), whereas the large-scale compartments analyzed here capture more stable, global segregation patterns. Understanding how these two regimes relate to one another remains an important open question.

      We have added the following statement in the Discussion acknowledging the scale-dependent nature of compartmentalization: “At the same time, recent ultra-high-resolution Hi-C studies [36,37] have revealed compartment-like features at much finer genomic scales, emphasizing that A/B compartmentalization is, to some extent, inherently scale-dependent. Understanding how these fine-scale, often transient micro-compartments relate to the more stable, large-scale segregation patterns described here will be an important direction for future studies.”

      Reviewer #2 (Public review):

      Summary:

      This manuscript by Choubani et al presents a technically strong analysis of A/B compartment dynamics across interphase using cell-cycle-resolved Hi-C. By combining the elegant Fucci-based staging system with in situ Hi-C, the authors achieve unusually fine temporal resolution across G1, S, and G2, particularly within the short G1 phase of mESCs. The central finding that A/B compartment strength increases abruptly at the G1/S transition, stabilizes during S phase, and subsequently weakens toward G2 challenges the prevailing view that compartmentalization strengthens monotonically throughout interphase. The authors further propose that this "compartment maturation" is triggered by S-phase entry but occurs independently of active DNA synthesis, and that it involves a consolidation and large-scale reorganization of A-compartment domains.

      Strengths:

      Overall, this is a thoughtfully executed study that will be of broad interest to the 3D genome community. The data are of high quality, and the analyses are extensive, albeit not completely novel. In particular, previous work (Nagano et al 2017 and Zhang et al 2019) has shown that compartments are re-established after mitosis and strengthened during early interphase, and single-cell Hi-C studies have reported changes in compartment association across S phase. In particular, Nagano et al show that DNA replication correlates with a build-up of compartments, similar to what is presented here, with the authors' conclusion that compartment strength peaks in early S. The idea that it weakens toward G2, rather than continuing to strengthen, appears to be novel and differs from the prevailing framing in the literature.

      We thank Reviewer #2 for their thoughtful assessment and critique. We address their specific concerns below.

      Weaknesses:

      That said, several aspects of the conceptual framing and interpretation would also benefit from further clarification, and the mechanistic interpretation of the reported compartment dynamics requires more careful positioning relative to established models of genome organization. Specific concerns are outlined below:

      (1) One of the major conclusions of the study is that compartment maturation does not require ongoing DNA replication. However, the interpretation would benefit from more precise wording. Thymidine arrest still permits licensing, replisome assembly, and other S-phase-associated chromatin changes upstream of bulk DNA synthesis. Therefore, their data, as presented, demonstrate independence from DNA synthesis per se, but not necessarily from the broader replication program. Please clarify this distinction in the text and interpretations throughout the manuscript.

      We thank the reviewer for this important distinction. We agree with their point and have never claimed that compartment maturation is independent of the broader replication program. That is why we carefully used the term "active DNA synthesis" rather than "replication" throughout the manuscript.

      However, we acknowledge that one sentence in the text was ambiguous. The original sentence read: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a pre-replicative state where replication had not yet initiated, although cell-cycle markers indicated entry into S-phase.”

      We have now revised it to: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a state where the replication program (including origin licensing, replisome assembly, and helicase activation) has been initiated, as indicated by cell-cycle markers, but ongoing DNA synthesis (elongation) is blocked. ”

      This clarifies that compartment maturation is independent of active DNA synthesis (elongation) but not necessarily independent of upstream replication-associated processes. The change has been made in the manuscript.

      (2) A major conceptual issue that is not addressed at all is the well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartmentalization. Numerous studies have shown that loss of cohesin or reduced loop extrusion leads to stronger compartment signals, whereas increased cohesin residence or enhanced extrusion weakens compartmentalization. Given this framework, an obvious alternative explanation for the authors' observations is that the abrupt increase in compartment strength at G1/S, and its decline toward G2, could reflect cell-cycle-dependent modulation of cohesin activity rather than a compartment-intrinsic "maturation" program.

      The manuscript does not explicitly consider this possibility, nor does it examine loop extrusion-related features (such as loop strength, insulation, or stripe patterns) across the same cell-cycle stages. Without discussing or analyzing this widely accepted model, it is difficult to distinguish whether the reported compartment dynamics represent a novel architectural mechanism or an indirect consequence of known changes in extrusion behavior during the cell cycle. I strongly encourage the authors to analyze their data to determine if they observe anti-correlated loop changes at the same time they observe compartment changes. Ideally, the authors would remove loop extrusion during interphase using well-established cohesin degrons available in mESCs and determine if the relative differences in compartment dynamics persist.

      We thank the reviewer for raising this interesting point. We agree that there is a well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartment strength in the literature.

      To test whether cell cycle compartment dynamics, particularly compartment maturation at the G1/S transition, could be explained by changes in loop extrusion, we analyzed insulation at RAD21/CTCF sites (mESC data from Hansen et al., eLife, 2017) across the cell cycle. During normal cycling, we indeed observed an anti-correlation: insulation dropped as compartment strength increased at the G1/S transition. However, in G1/S-arrested cells, insulation did not drop compared to late G1 (it even slightly increased) even though compartment maturation still occurred, indicating that the two processes can be uncoupled. This is consistent with other studies showing that loop extrusion and compartment dynamics are driven by independent mechanisms (Nora et al., Cell, 2017; Zhang et al., Nat Commun, 2021), although we cannot fully rule out some contribution from loop extrusion dynamics without direct cohesin degron experiments.

      We have added a new Results section describing these findings titled “Compartment maturation is independent of cohesin-mediated loop extrusion”, including new Figure panels 3H, I, and Figure S7.

      (3) The proposed "peninsula-like" A-domain structures are inferred from ensemble Hi-C data and polymer modeling, rather than directly observed physical conformations. That is, single-cell imaging data clearly have shown that Hi-C (especially ensemble Hi-C) cannot uniquely specify physical conformations and that different underlying structures can produce similar contact patterns. The "peninsula" language, as written, risks being interpreted as a literal structural model rather than a conceptual visualization. Instead of risking this as just another nuanced Hi-C feature in the field, the authors could strengthen the manuscript by either (i) explicitly framing the peninsula model as a heuristic description of contact redistribution rather than a definitive physical architecture, or (ii) discussing alternative structural scenarios that could give rise to similar Hi-C patterns. Clarifying this distinction would improve the rigor and help readers better understand what aspects of A-compartment consolidation are directly supported by the data versus model-based extrapolations. For example, it would be useful to clarify whether the observed increase in long-range A-A contacts reflects spatial extension of internal A regions, changes in loop extrusion dynamics, increased compartment mixing within the A state, or population-averaged heterogeneity across alleles.

      We thank the reviewer for this important clarification. We agree that the "peninsula" model should be framed as a heuristic description. As detailed in our response to Reviewer #1 (see above), we have added a disclaimer to the manuscript and provided orthogonal DNA FISH support for physical extension of A-domains during S phase. We have also ensured that the language emphasizes the conceptual nature of the model.

      (4) The extension of the analysis to additional cell types using HiRES single-cell data is a valuable addition and supports the idea that compartment maturation is not unique to mESCs. However, the limitations of these data, in particular, the limited phase resolution, in addition to the pseudo-bulk aggregation and variable coverage, should be emphasized more clearly in the main text. Framing these results as evidence for conservation in principle, rather than definitive proof of identical dynamics across tissues, would be a more appropriate framing.

      We agree with the reviewer. We have already explicitly acknowledged the limited temporal resolution and variable coverage of the HiRES dataset in the main text. To better reflect its supporting role, we have moved the HiRES figure (previously Fig. 4) to Fig. S10 and merged the corresponding results section with the previous one titled: “Formation of a consolidated A compartment in S-phase”.

      We have also revised the language to avoid overstatement. The original conclusion read: “Together, these findings strongly indicate that compartment maturation and the accompanying A compartment consolidation represent a robust and universally observed feature across different developmental contexts.”

      This has been changed to: “Together, these findings support the notion that compartment maturation and the accompanying A-compartment consolidation are not unique to mESCs and may represent a broadly conserved feature of mammalian chromatin organization.”

      Similarly, the abstract has been adjusted from: “Moreover, compartment maturation was not limited to mESCs but was also observed across different developmental contexts in mice.” to: “Moreover, compartment maturation was not limited to mESCs but was also evident across different developmental contexts in mice.”

      These changes frame the results as evidence for conservation in principle rather than definitive proof of identical dynamics across tissues.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please address the minor points in the public review.

      In addition, on page 7, line 285: "In contrast, interactions showed minimal change across all distances though interphase". Do the authors mean "In contrast, B-B interactions..."?

      We thank the reviewer for catching this. The sentence has been corrected.

    1. eLife Assessment

      This important study identifies PRRT2 as an auxiliary regulator of Nav channel slow inactivation in vitro and in vivo, proposing that PRRT2 facilitates entry into, and delays recovery from, the slow-inactivated state. The revised manuscript has been substantially strengthened, providing compelling evidence that PRRT2 is relevant to normal brain physiology and disease pathophysiology, providing a mechanistic link between PRRT2 mutations and episodic neurological phenotypes. Overall, this study will be of interest to ion channel biophysicists and neurophysiologists, particularly those studying channelopathies.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Lu and colleagues demonstrate convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      Comments on revised version.

      The manuscript by Lu and colleagues has been revised sufficiently to address all my prior concerns.

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

    3. Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".<br /> PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last (20th) compound APs in panels B and C.

    4. Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      (4) The mechanistic separation between trafficking of PRRT2 and its gating effects is not clearly resolved.

      (5) Additional studies with Nav1.6 should be carried out.

      Comments on revised version.

      These comments have been addressed in the revised version.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Lu and colleagues demonstrates convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously, including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      We thank the reviewer for these positive comments and for the thoughtful evaluation of our work.

      Weaknesses:

      There are a few missing experiments and one place where data are over-interpreted.

      (1) An in vitro study of Nav1.6 is conspicuously absent. In addition to being a major brain Na channel, Nav1.6 is predominant in cerebellar Purkinje neurons, which the authors note lack PRRT2 expression. They speculate that the absence of PRRT2 in these neurons facilitates the high firing rate. This hypothesis would be strengthened if PRRT2 also enhanced slow inactivation of Nav1.6. If a stable Nav1.6 cell were not available, then simple transient co-transfection experiments would suffice.

      We thank the reviewer for raising this point. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms.

      We have now performed new heterologous expression experiments to test whether PRRT2 modulates Nav1.6 slow inactivation. Consistent with our findings for other Nav isoforms, PRRT2 significantly enhances the slow inactivation of Nav1.6. We have incorporated these data into the revised Results and Figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) To further demonstrate the physiological impact of enhanced slow inactivation, the authors should consider a simple experiment in the stable cell line experiments (Figure 1) to test pulse frequency dependence of peak Na current. One would predict that PRRT2 expression will potentiate 'run down' of the channels, and this finding would be complementary to the biophysical data.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we performed a pulse-train protocol in the stable Nav1.2 cell line and quantified the use-dependent attenuation (“run-down”) of peak sodium current across successive depolarizations (Figure 1-figure supplement 1C). Compared with control cells, PRRT2-expressing cells exhibited a larger decline in peak current during trains, indicating greater reduction in channel availability during repetitive depolarizations (Figure 1-figure supplement 1C). This pattern is consistent with our observations above showing that PRRT2 enhances Nav channel slow inactivation. These new data have been incorporated into the revised manuscript. Please refer to Page 5, Lines 133-140; Figure 1-figure supplement 1C.

      (3) The study of one K channel is limited, and the conclusion from these experiments represents an over-interpretation. I suggest removing these data unless many more K channels (ideally with measurable proxies for slow inactivation) were tested. These data do not contribute much to the story.

      We agree with the reviewer’s assessment. To avoid over-interpretation and to maintain focus on PRRT2-dependent regulation of Nav channel slow inactivation, we have removed the potassium channel dataset and the associated conclusions from the revised manuscript.

      (4) In Figure 2, the authors should confirm that protein is indeed expressed in cells expressing each truncated PRRT2 construct. Absent expression should be ruled out as an explanation for the enhancement of slow inactivation.

      We thank the reviewer’s concern regarding expression of the truncated PRRT2 constructs in the Nav1.2 stable cell line, particularly PRRT2(1-266), which shows little effect on slow inactivation of Nav1.2 channels. In the revised manuscript, we conducted western blot to verify expression of the PRRT2(1-266)-HA construct in the Nav1.2 stable cell line. We have added these results to the revised manuscript, please refer to Page 6, Lines 171-173; Figure 2-figure supplement 1A and B.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is primarily expressed in the nervous system and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 directly interacts with Nav1.2 and Nav1.6, modulating channel properties and neuronal excitability.

      In this study, Lu et al. reported that PRRT2 is a physiological regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect can be replicated by the C-terminal region (256-346) of PRRT2, and is highly conserved across species from zebrafish, mouse, to human PRRT2. TRARG1 and TMEM233, the other two DspB family members, showed similar effects on Nav1.2 slow inactivation. Co-IP data confirms the interaction between Nav channels and PRRT2. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared to WT mice.

      Strengths:

      (1) This study is well designed, and data support the conclusion that PRRT2 is a potent regulator of slow inactivation of Nav channels.

      (2) This study reveals similar effects on Nav1.2 slow inactivation by PRRT2, TMEM233, and TRARG1, indicating a common regulation of Nav channels by DspB family members (Supplemental Figure 2). A recent study has shown that TMEM233 is essential for ExTxA (a plant toxin)-mediated inhibition on fast inactivation of Nav channels; and PRRT2 and TRARG1 could replicate this effect (Jami S, et al. Nat Commun 2023). It is possible that all three DspB members regulate Nav channel properties through the same mechanism, and exploring molecules that target PRRT2/TRARG1/TMEM233 might be a novel strategy for developing new treatments of DspB-related neurological diseases.

      We thank the reviewer for careful evaluation and insightful suggestions.

      Weaknesses:

      (1) Previously, the authors have reported that PRRT2 reduces Nav1.2 current density and alters biophysical properties of both Nav1.2 and Nav1.6 channels, including enhanced steady-state inactivation, slower recovery, and stronger use-dependent inhibition (Lu B, et al. Cell Rep 2021, Fig 3 & S5). All those changes are expected to alter neuronal excitability and should be discussed.

      We thank the reviewer for this suggestion. Although the present study focuses on PRRT2-dependent regulation of slow inactivation, we agree that PRRT2 may influence excitability through additional Nav-dependent mechanisms, including reduced current density and shifts in the voltage dependence of channel inactivation (Fruscione et al., 2018; Lu et al., 2021; Valente et al., 2023). Notably, because PRRT2 facilitates entry of Nav channels into slow-inactivated states both from closed states and from open states during prolonged depolarization, some of these previously reported effects may partly reflect enhanced slow inactivation and the resulting reduction in Nav channel availability. We have expanded the Discussion to integrate these prior findings and to clarify that these additional PRRT2-dependent effects may converge to shape neuronal excitability. Please refer to Page 16, Lines 445-452.

      (2) In this study, the fast inactivation kinetics was examined by a single stimulus at 0 mV, which may not be sufficient for the conclusion. Inactivation kinetics at more voltage potentials should be added.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we expanded our analysis of Nav1.2 fast-inactivation kinetics to include a range of test potentials (-20, -10, 0, +10, +20 and +30 mV) in the presence and absence of PRRT2. These experiments showed that PRRT2 expression did not significantly affect Nav1.2 fast-inactivation kinetics under these conditions. We have incorporated these new results into the revised manuscript. Please refer to Page 4, Lines 100-103; Figure 1C.

      (3) It is a little surprising that there is no difference in Nav1.2 current density in axon-blebs between WT and Prrt2-mutant mice (Figure 7B). PRRT2 significantly shifts steady-state slow inactivation curve to hyperpolarizing direction, at -70 mV, nearly 70% of Nav1.2 channels are inactivated by slow inactivation in cells expressing PRRT2 when compared to less than 10% in cells expressing GFP (Figure supplement 1B); with a holding potential of -70 mV, I would expect that most of Nav channels are inactivated in axon-blebs from WT mice but not in axon-blebs from Prrt2-mutant mice, and therefore sodium current density should be different in Figure 7B, which was not. Any explanation?

      We thank the reviewer for raising this point. In our axonal bleb recordings, although the holding potential was -70 mV, sodium current density was measured after a hyperpolarizing pre-pulse to -110 mV, which was applied before the test depolarization to relieve inactivation as much as possible (as described in the Methods). Therefore, the current density measurement in Figure 7B reflects the available current after this recovery step, rather than the steady-state availability at -70 mV. The lack of a difference in Figure 7B does not contradict the PRRT2-dependent shift in steady-state slow inactivation. In the revised manuscript, we have clarified this point explicitly in the Results and figure legend to avoid confusion. Please refer to Page 10, Lines 294-295.

      (4) Besides Nav channels, PRRT2 has been shown to act on Cav2.1 channels as well as molecules involved in neurotransmitter release, which may also contribute to abnormal neuronal activity in Prrt2-mutant mice. These should be mentioned when discussing PRRT2's role in neuronal resilience.

      We thank the reviewer for this suggestion. In addition to the Nav-dependent mechanisms, previous studies have shown that PRRT2 also regulates synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018) and presynaptic surface expression of Cav2.1 channels (Ferrante et al., 2021). These effects are also expected to influence neurotransmitter release and, consequently, neuronal and network excitability. In the revised manuscript, we have expanded the Discussion to acknowledge that these additional PRRT2-dependent mechanisms may also contribute to cortical resilience. Please refer to Page 16, Lines 452-457.

      Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      We thank the reviewer for this positive evaluation of our work and for the constructive comments.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav channel interface—through approaches such as targeted mutagenesis, crosslinking, and structural determination—will be required to define the binding interface and establish the molecular basis of gating modulation. Please refer to Page 16, Lines 465-468.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      We agree with the reviewer. Impaired slow inactivation in Prrt2-mutant mice is one plausible contributor to reduced cortical resilience. PRRT2 has also been reported to regulate surface exposure of Nav and Cav2.1 channels (Ferrante et al., 2021), as well as neuronal synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018). Each of these PRRT2-associated processes could influence cortical excitability in vivo. We have therefore expanded the Discussion to clarify that the cortical phenotype likely reflects the combined contribution of multiple PRRT2-dependent mechanisms, rather than an isolated defect in slow inactivation alone. Please refer to Page 16, Lines 446-458.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      We thank the review for this comment regarding physiological relevance. In the revised manuscript, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Please refer to Page 14 and 15, Lines 414-416; Lines 429-430.

      (4) The mechanistic separation between the trafficking effect of PRRT2 and its gating effects is not clearly resolved.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. Previous studies in heterologous overexpression systems have shown that PRRT2 can influence Nav channel trafficking and surface expression, raising the possibility that the observed effects on slow inactivation regulation might be secondary to altered channel abundance or localization. However, slow inactivation develops on a timescale of tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer intervals (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). These distinct temporal profiles argue against trafficking as the primary basis for the effects of PRRT2 on Nav channel slow inactivation described here, although direct quantification of dynamic changes in Nav channel surface expression will be required to fully exclude such a contribution (Liu et al., 2022; Tyagi et al., 2025). We have incorporated this point into the Discussion section. Please refer to Pages 13, Lines 378-388.

      (5) Additional studies with Nav1.6 should be carried out.

      We thank the reviewer for this suggestion. We have performed experiments to directly examine the effects of PRRT2 on Nav1.6 slow inactivation and incorporated these new data into the revised Results and figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for future experiments (not for this paper)

      (1) Exploit the lower protein expression in V5-PRRT2 mice to examine the effects of a hypomorphic allele.

      We thank the reviewer for this insightful suggestion. We note that the V5 epitope knock-in reduced PRRT2 protein expression, which may functionally resemble a hypomorphic allele. Accordingly, in addition to its utility for biochemical experiments (e.g., co-immunoprecipitation), this line could serve as a genetic tool to interrogate PRRT2 dose-dependent effects in vivo. We have added this point to the revised manuscript, please refer to Page 9, Lines 265-267.

      (2) Examine disease-causing PRRT2 mutations.

      We thank the reviewer for this constructive suggestion. Testing disease-associated PRRT2 variants for their ability to regulate Nav channel slow inactivation would be an important next step to strengthen the disease relevance of the mechanism proposed here. Moreover, identifying missense variants that selectively disrupt slow-inactivation regulation could help pinpoint residues that are critical for PRRT2-Nav functional coupling and thereby inform future structure-function studies. We plan to pursue this direction in follow-up work.

      (3) Investigate spreading depolarization in PRRT2-deficient mice.

      We thank the reviewer for this suggestion. Although we have shown that PRRT2 deficiency facilitates spreading depolarization in the cerebellum, whether PRRT2 exerts similar control over spreading depolarization susceptibility in the cerebral cortex remains to be determined. We plan to address this in an independent study and to test how cortical spreading depolarization relates to other PRRT2-associated neurological disorders.

      Reviewer #2 (Recommendations for the authors):

      This study is, in general, well executed, and the manuscript is well written. However, I do have some questions.

      (1) The authors' previous works have shown that PRRT2 regulates both Nav1.2 and Nav1.6, considering the wide expression Nav1.6 in CNS and its role in neuronal activity, what makes the authors not include Nav1.6 in this study?

      We thank the reviewer for raising this question. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms. In response to reviewers’ concern, we have now performed new experiments to directly examine the effect of PRRT2 on Nav1.6 slow inactivation. These results have been incorporated into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) Please explain why you chose 0 mV rather than -70 mV (closer to membrane potential) in the slow inactivation protocol.

      We thank the reviewer for raising this question. Nav channels can enter into slow inactivation from both resting/closed states and activated/open states. In our steady-state slow-inactivation assays, we found that PRRT2 enhances Nav1.2 slow inactivation under both conditions (Figure 1-figure supplement 1A and B). In whole-cell recordings, Nav1.2 channels typically begin to activate at command voltages more depolarized than approximately -60 mV. Accordingly, a conditioning voltage of -70 mV predominantly probes entry into slow inactivation from closed states, whereas 0 mV drives channel activation and more effectively induces slow inactivation. We therefore chose 0 mV as the primary conditioning potential because it is widely used in conventional slow inactivation protocols and induces slow inactivation more robustly than conditioning voltages at -70 mV. We have added this explanation in Methods section of revised manuscript, please refer to Page 20, Lines 569-571.

      (3) The authors mentioned that the insertion of V5 markedly reduced the PRRT2 protein level; thus, Prrt2-V5 knock-in mice could be considered as PRRT2 knock-down mice. Is there any noticeable difference in phenotype between Prrt2-V5 knock-in mouse and Prrt2-mutant mouse? In other words, is PRRT2 knockdown sufficient to affect neuronal excitability, or is a complete PRRT2 ablation required?

      We thank the reviewer for raising this concern regarding the functional consequences of reduced PRRT2 expression in the Prrt2-V5 knock-in mice. Given that PRRT2 protein levels are markedly reduced in this line, and that cerebellar stimulation-induced dystonia is a characteristic phenotype of PRRT2 deficiency, we tested whether Prrt2-V5 knock-in mice also exhibit this phenotype. We found that electrical stimulation of the cerebellar cortex induced dystonia-like attacks in a subset of Prrt2-V5 knock-in mice. These dystonic behaviors resembled those previously observed in Prrt2-mutant mice, whereas no such behaviors were induced in wild-type mice (Figure 6-figure supplement 1). These findings indicate that a substantial reduction of PRRT2 expression (approximately 80%) is sufficient to impair neuronal function and elicit a disease-relevant phenotype in a subset of animals, supporting the interpretation that the V5 knock-in allele is hypomorphic. We have incorporated these results into the revised manuscript, please refer to Page 9, Lines 265-267; Figure 6-figure supplement 1.

      (4) In Discussion (Page 13, lines 358-361), the authors mentioned a putative interaction between PRRT2 and the Nav channel by modeling, while there is no related data. Please either add modeling data or remove those sentences.

      We thank the reviewer for this suggestion. To avoid over-interpretation, we have removed the statements regarding the AlphaFold-based interaction model from the revised manuscript. We agree that the interaction interface remains to be demonstrated experimentally, and we now discuss this point in the Limitations section. Please refer to Page 16, Lines 465-468.

      (5) Typo: Page 14, line 399, "TMEM232" should be "TMEM233".

      We thank the reviewer for pointing out this typo. We have corrected it in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Mechanistic depth: While the functional data show altered slow-inactivation kinetics, the mechanistic explanation remains superficial. The AlphaFold-based prediction of PRRT2 interaction with DIV-S3 is speculative. The authors should clarify their illustrative rather than evidential intent and avoid over-interpretation.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav interface, including targeted mutagenesis, crosslinking, and structural determination, will be necessary to elucidate the molecular basis of this interaction and its effect on channel gating. Please refer to Page 16, Lines 465-468.

      (2) Separation of trafficking vs. gating effects: Previous studies showed PRRT2 influences Nav trafficking and surface expression. Here, surface expression changes are not systematically quantified. Such an analysis would strengthen the argument that gating effects are not secondary to altered channel abundance or localization.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. We agree that direct analysis of Nav channel surface localization during prolonged depolarization and hyperpolarization would provide stronger evidence to distinguish gating effects from trafficking-dependent mechanisms. However, such experiments are technically challenging in this context: conventional surface biotinylation assays do not provide the temporal resolution required for these rapid protocols, and live-cell imaging approaches to monitor dynamic changes in Nav channel surface expression during slow-inactivation paradigms have not yet been established in our laboratory.

      Although PRRT2 has been reported to regulate Nav channel surface expression in heterologous systems, we consider it unlikely that trafficking is the major determinant of the slow-inactivation effects described here. Slow-inactivation develops on a timescale ranging from tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer timescales (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). We have expanded the Discussion in a revised manuscript. Please refer to Pages 13, Lines 378-388.

      (3) Isoform generalization: Data on other Nav channel subtypes are presented as evidence of a conserved mechanism. However, given tissue-specific expression of PRRT2, these findings may be of limited in vivo relevance. At the very least, additional studies with Nav1.6 should be carried out.

      We thank the review for this suggestion. In response, we conducted new experiments to examine the effect of PRRT2 on Nav1.6 slow inactivation. These results show that PRRT2 promotes entry of Nav1.6 channels into slow-inactivated states and delays their recovery, consistent with its effects on the other Nav isoforms examined in this study. We have incorporated these new data into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      Furthermore, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Pages 14 and 15, Lines 414-416 and 429-430.

      (4) In vivo functional link: The EEG after-discharge threshold assay suggests decreased cortical resilience, but causality between slow-inactivation impairment and hyperexcitability remains indirect. Complementary in vivo recordings would strengthen the physiological link.

      We thank the reviewer for this helpful suggestion. To further link impaired slow-inactivation to the hyperexcitability, we applied a repetitive stimulation protocol in corpus callosum slices, a white-matter region of brain enriched in both PRRT2 and Nav channels. During high-frequency stimulation (e.g., 20 Hz), the amplitude of the compound action potential progressively decreased over the course of the stimulus train. This phenomenon, often referred to as adaptation, reflects activity-dependent reduction in Nav channel availability (Fleidervish et al., 1996; Mickus et al., 1999; Kim et al., 2012). Compared with wild-type mice, Prrt2-mutant mice exhibited less adaptation during high-frequency stimulation, consistent with impaired slow inactivation during repetitive activity, which may contribute to hyperexcitability (Figure 7-figure supplement 2). We have added these results to the revised manuscript. Please refer to Pages 11, Lines 311-322; Figure 7-figure supplement 2.

      (5) Structural interaction: It remains unclear whether PRRT2 binds the α-subunit directly or through accessory proteins. Crosslinking or detergent-solubilization controls of different stringencies could clarify this.

      We thank the reviewer for raising this important issue. We agree that our co-immunoprecipitation data do not distinguish whether PRRT2 associates with the Nav channel α-subunit directly or through other components of the protein complex. To avoid over-interpretation, we have revised the relevant text in the manuscript to remove any implication of direct binding and now describe the result as an association between PRRT2 and Nav channels.

      We have also expanded the Limitations section to note that additional experiments, such as crosslinking and structural studies, will be required to define the interaction interface between PRRT2 and Nav channels. Please refer to Page 16, Lines 465-468.

      (6) Comparisons to other regulators: The paper positions PRRT2 as distinct from FHFs and β-subunits. The data support this, but the discussion could more critically assess whether PRRT2 acts by stabilizing a pore-based inactivated conformation, as suggested for other slow-inactivation modulators.

      We thank the reviewer for this insightful suggestion. At present, relatively few modulators have been characterized in detail with respect to their effects on Nav channel slow-inactivation kinetics. Moreover, even for compounds such as lacosamide, which has been proposed to act as a slow-inactivation modulator, the underlying mechanism remains under debate (Errington et al., 2008; Jo and Bean, 2017). Therefore, in the revised manuscript, we discussed the possible mechanism of PRRT2 in the context of current models of Nav channel slow inactivation.

      Previous studies suggest that entry into the slow-inactivated state involves at least two coupled processes: conformational changes in the voltage-sensing domains and structural rearrangements in the pore region, including the selectivity filter and intracellular activation gate (Catterall et al., 2024; Silva, 2014). During prolonged depolarization, voltage sensors become stabilized in the up-state, while the pore undergoes progressive rearrangements associated with slow inactivation (Balser et al., 1996; Vilin et al., 1999). Thus, mechanisms that further stabilize voltage sensors in the up-state and/or facilitate pore-based inactivated conformations could enhance slow inactivation.

      Within this framework, PRRT2 may enhance slow inactivation by facilitating one or both of these processes, although direct evidence is still lacking. We have incorporated this discussion in relative section of revised manuscript. Please refer to Page 14, Lines 389-404.

      Response references:

      Jo S, Bean BP. Lacosamide Inhibition of Nav1.7 Voltage-Gated Sodium Channels: Slow Binding to Fast-Inactivated States. Mol Pharmacol. 2017 Apr;91(4):277-286.

      Errington AC, Stöhr T, Heers C, Lees G. The investigational anticonvulsant lacosamide selectively enhances slow inactivation of voltage-gated sodium channels. Mol Pharmacol. 2008 Jan;73(1):157-69.

      (7) Behavioral/clinical link: Given the strong human genetics background of PRRT2 disorders, a brief analysis or reference to electrophysiological phenotypes in patient neurons would contextualize the cortical findings.

      We thank the reviewer for this suggestion. Previous studies showed that iPSC-derived excitatory neurons from a patient carrying a homozygous PRRT2 mutation exhibited increased sodium currents and neuronal hyperexcitability (Fruscione et al., 2018). Given that slow inactivation regulates Nav channel availability and thereby influences neuronal excitability, these electrophysiological abnormalities in patient-derived neurons may, at least in part, reflect impaired PRRT2-dependent regulation of Nav channel slow inactivation. We have added this point to the relative section of the revised manuscript. Please refer to Pages 15, Lines 432-437.

      Minor comments

      (1) Figures should include statistical sample sizes (n) and ideally overlay data points rather than only means {plus minus} SEM.

      We thank the reviewer for this suggestion. In the revised manuscript, we present both individual data points and mean ± SEM in the column graphs. For the line graphs, individual data points were not overlaid because of space and readability constraints, and these panels therefore display mean ± SEM only. Sample sizes for each group are provided in the corresponding figure legends.

      (2) The AlphaFold model should be provided as a supplementary figure with confidence scores indicated.

      We thank the reviewer for this suggestion. However, because the predicted Nav1.2-PRRT2 interaction interface has not yet been experimentally validated in our study, we chose to remove the AlphaFold-based model from the revised manuscript to avoid over-interpretation.

      (3) Clarify whether TTX sensitivity was verified in the axonal bleb preparation.

      We thank the reviewer for raising this point. We verified the identity of the sodium currents in the axonal bleb preparation by their sensitivity to TTX, and this information has now been added to Figure 7A in the revised manuscript. Please refer to Page 10, Line 290; Figure 7A.

    1. eLife Assessment

      This important study investigates how surface stickiness shapes whisker mechanics and peripheral neural responses during active touch. The biomechanical evidence that surface stickiness alters whisker mechanics and stick-slip dynamics is compelling, supported by a large and high-quality 3D dataset, while the electrophysiological evidence is solid but limited by a small sample size and insufficient validation of the sticky stimuli. The work will be of broad interest to sensory neuroscientists studying active touch.

    2. Reviewer #1 (Public review):

      Summary:

      This study offers a careful and technically strong look at how surface stickiness changes whisker-surface interactions and how that information reaches peripheral sensory neurons. The authors use 3D whisker tracking to capture bending, twisting, rolling, and tip motion during contact with surfaces that differ in stickiness, coarseness, and position. They show that sticky surfaces, especially silicone, broaden the range of whisker deformation, produce stronger but less frequent stick-slip events, and change firing rates in some trigeminal ganglion neurons. Overall, the study is valuable because it goes beyond standard 2D tracking and shows that out-of-plane motion and roll are important for understanding how whiskers encode texture.

      Strengths:

      The study is technically strong and well motivated. Its main strength is the use of 3D whisker tracking to show that surface stickiness affects whisker deformation in ways that standard 2D tracking would miss, including torsion, roll, out-of-plane motion, and stick-slip dynamics. The authors also connect these mechanical effects to TG activity, providing evidence that stickiness information is available in peripheral sensory responses. Overall, the work expands the study of whisker-based texture sensing beyond coarseness and provides a richer biomechanical framework for understanding tactile encoding.

      Weaknesses:

      The main weakness is that stickiness is not formally defined early in the manuscript, even though it is the central experimental variable. Several methodological choices also need clearer justification or validation, including the use of 2D measures as comparators for torsion and roll, the thresholds used for stick-slip detection, the degree-5 polynomial fit, the reference ROI, and aspects of the 3D surface reconstruction. The neural evidence should also be interpreted cautiously because the TG sample is small, only a subset of units discriminated silicone, and the correlation between strain sensitivity and silicone discrimination is suggestive rather than definitive.

    3. Reviewer #2 (Public review):

      The authors explore the sensation of stickiness from the point of view of whisker exploration and encoding in the trigeminal ganglion. In doing so, they develop methods for 3D whisker tracking to describe stick-specific parameters such as stick-slip rates and strain. Overall, the methods are strong, and the authors present the results appropriately. Overall, I think exploration of the sensation of stickiness is a great question.

      My main criticism is in relation to the chosen stimuli, and I wonder whether the authors may have room to explore more naturally sticky materials and what this may mean for the animal.

      (1) Chosen stimuli for stickiness:

      Four different materials are used, with the aim of presenting animals with graded measures of stickiness. The results show that silicone stands out against the others; it's less clear whether the intermediate textures (Delrin and resin) may be truly intermediate in stickiness.

      I wonder if the stimuli chosen were truly representative of the aim of providing a gradient of stickiness. Did the materials differ in other features, such as surface temperature, texture, etc., which could explain some results? The authors discuss this in terms of coefficients of friction and how these estimates are not quantified in relation to whiskers themselves.

      Measures of stick-slip and strain with silicone vs other materials make intuitive sense. Could the authors add additional naturally sticky stimuli to exemplify the results? For example, adhesive, glue, or a sugary substance.

      (2) Tracking methods and quantification:

      The 3D tracking methods, which incorporate whisker twists, strain, and other fine features of whisker exploration, present an advance in terms of analysis of how whiskers may explore more complex, natural features of environments. The analyses and quantifications are all solid and robust. The technical approaches are well-prepared to take the work a step further in terms of stimulus choice.

      (3) Peripheral coding of stickiness:

      The authors report that some units respond preferentially to whisking on silicone and that this has to do with strain on the whisker. Is there a possibility to understand the nature or anatomy of these units and why they might be preferential for the sticky sensation? Can the location in the follicle be assigned? And/or would the methodology allow for assignment of where the specifically sticky-tuned units project centrally?

      (4) Relationship to natural stimuli:

      A piece missing from the paper is more discussion and exploration of why stickiness may be important for sensory coding, as well as potentially more naturally sticky stimuli. One could imagine that a mouse navigating the world could find stickiness attractive, if it were a source of sweet food, for example, or it could potentially be a sensation the animal prefers to avoid. Stickiness could also indicate contamination or a sticky trap, to be avoided. If the authors are able to add naturally sticky stimuli, the whisker exploration and encoding could potentially provide further cues towards the valence of stickiness for mice.

    4. Reviewer #3 (Public review):

      This paper tackles an underexplored dimension of whisker-based texture sensing: while surface coarseness encoding has been extensively characterized in rodents, the mechanical and neural basis for stickiness sensing has not previously been examined. The authors make two intertwined contributions that together represent a substantial advance: a methodological one - a 3D whisker tracking pipeline operating at 4000 fps, capable of capturing torsion, roll, and out-of-plane whisker motion - and a scientific one - a first characterization of how whisker mechanics and primary trigeminal afferent responses differ between surfaces of high and low stickiness. The work is technically solid, the dataset is large, and the question is well motivated both by the multidimensional nature of tactile texture perception and by the practical advantages of the whisker system for studying touch mechanics.

      Strengths.:

      The 3D tracking system is a timely advance over existing tools, particularly in its handling of non-planar whisker shapes and the full automation required for the sub-millisecond resolution needed to detect stick-slip events. The mechanical dataset is extensive. The finding that whisking against silicone expands the sampled whisker strain space and produces stronger but less frequent stick-slip events is clearly demonstrated and internally consistent with the proposed mechanism of greater strain accumulation before frictional release - a physically intuitive result. The open release of the tracking code considerably increases the value of this work to the broader community.

      Weaknesses:

      A few aspects of the paper, if sharpened, would considerably strengthen the evidence and the clarity of the conclusions.

      The central claim - that "stickiness information is available to the whisker system" - does not capture the precision of what the paper demonstrates. As stated, the finding is close to guaranteed: any variation in surface friction will produce some change in whisker mechanics, so the presence of mechanical differences between materials is expected rather than surprising. The more valuable question the paper is well positioned to answer is which specific dimensions of the whisker mechanical response are most informative about surface stickiness. The paper reports effects on strain distribution breadth, stick-slip amplitude, and stick-slip rate, but does not synthesize which of these - or which sub-dimensions (bending, twisting, or rolling) - carry the most discriminating information. Identifying the salient dimensions of the mechanical response and relating them to the proposed frictional mechanism would sharpen the paper's conclusions substantially.

      A related but distinct limitation is the absence of direct force measurements during whisker-surface contact. The authors acknowledge this openly, and I recognize it is not easily remedied within the current experimental setup. It does, however, constrain interpretation: without knowing the actual forces generated at the whisker-surface interface, the assumed stickiness ordering of the tested materials cannot be validated, and - importantly - the relative contribution of surface friction and material compliance to the observed mechanical differences cannot be determined. This is an important direction for future work in this area.

      The paper argues carefully that 2D tracking is insufficient for capturing the full mechanical picture of whisker-surface interactions, and the figure currently in the supplementary material (Figure S2) makes this case convincingly through multiple analyses. This argument is the core justification for the paper's methodological contribution and deserves a place in the main manuscript. Furthermore, while the mechanical case for 3D over 2D tracking is well made, it has not yet been tested at the neural level: the regression model used to predict neural firing incorporates 3D variables, but its performance is not compared against an equivalent model restricted to 2D variables. Such a comparison would directly demonstrate whether torsion and roll - the signals inaccessible to 2D tracking - carry neural predictive value, and would elegantly unite the paper's methodological and scientific contributions.

      Finally, the three-dimensional plots in Figure 3 are the paper's primary representation of its main mechanical result, and there is a real opportunity to make them considerably more informative. The whisker deformation probability distributions (panel B) are rendered in 3D from a single viewing angle, making it difficult to assess the shape or anisotropy of the distributions - and in particular to see which dimensions expand most for silicone relative to the other materials. This is precisely the information needed to identify the most salient dimensions of the stickiness signal, and two-dimensional representations would make it directly readable.

    1. eLife Assessment

      This study presents a useful compendium of triangulated single-cell eQTLs, Mendelian randomisation and colocalization of genetic signals in prostate cancer datasets. Biological interpretation in the context of the aging prostate gland, the tumour microenvironment and immune cell specificity is incomplete, so this study is a starting point for further study, and would require validation of the resulting putative causal genes.

    2. Reviewer #1 (Public review):

      Summary:

      Using Mendelian randomisation on available GWAS data, the investigators identified eGenes associated with prostate cancer and applied the data to define relevant immune cell types involved. Additional analysis was performed to explore potential candidate targets and agents from licensed medicines.

      This is an interesting approach as the investigators have expertise in other research fields, applied here to prostate cancers. The use of three different datasets is significant, and the approach to further analyse implicated eGenes in drug target analysis is relevant and timely.

      A particular strength is taking putative genes from Mendelian randomisation analysis to target and potential drug agents.

      Some aspects of the study would need to be clarified to enable interpretation of the findings in the context of the prostate gland and prostate cancers: expanding the descriptions of the supporting Supplementary Data and Tables, explanations of the analysis for the general reader, and clarification of the selection of eGenes (Figure 5).

    3. Reviewer #2 (Public review):

      Summary:

      This study integrates bulk and single-cell transcriptomic-derived eQTLs from two separate consortia (PRACTICAL and Finngen) to identify immune-cell-specific therapeutic targets in prostate cancer. Mendelian randomization and Bayesian colocalization have been used to produce druggable eGene modules through STRING and DrugBank.

      This is an interesting study that is attempting to address risk-associated, immune-specific transcriptomic repertoires in prostate cancer. It is knitting together concepts of drug repurposing and prostate cancer immunogenicity. This is an entirely computational study, which would benefit from some wet lab experimental validation.

      It is very tricky to attribute cell-type-specific responses, especially when the majority of genes involved represent cytoskeletal or stress responses, which are ubiquitous throughout the prostate microenvironment. This point is relevant for the drug repurposing section: if these drugs are targeting immune cell-specific repertoires, what would the response be of the entire environment? It would be useful to contextualize the validity of each proposed therapy in a specific prostate cancer context and the involvement of AR antagonism or radiotherapy.

      Strengths and limitations of this study:

      Strengths:

      This is a scientifically interesting and potentially impactful study, particularly in its attempt to integrate immune-cell-specific transcriptomics, causal inference, and drug repurposing in prostate cancer. The methodology is well described, and the data (albeit limited) are well analyzed.

      Limitations:

      The central weakness is the overstatement of the conclusions regarding immune-cell-specific causality, without sufficiently contextualizing the biological meaning of the findings.

      Highlighted genes, such as LMNA, XBP1, histone-related genes, and stress-response markers, are ubiquitous regulators involved in fundamental cellular processes, including ageing, unfolded protein response (UPR), integrated stress response (ISR), chromatin remodeling, proliferation, and metabolism. It is unclear whether these signatures truly represent immune mechanisms, or instead reflect broader inflammatory and age-associated biology expected within an ageing glandular organ such as the prostate.

      Immune cell identity alone may not be sufficient to infer biological relevance because immune state characterization (e.g., exhausted versus functional T cells, or distinct macrophage/myeloid phenotypes) is largely absent from the current analysis. The assertion that specific immune populations are correlated with prostate cancer susceptibility is probably an overstatement unless the nature of these cells can also be characterized.

      The interpretation of "causal variants" is not always specified, i.e., what phenotype is being associated: prostate cancer susceptibility, recurrence, progression, or treatment response (e.g. is there direct causality from immune-cell variants to prostate cancer?).

      Overall, there is a need for stronger biological and translational contextualization: how do the identified pathways relate to ageing-associated inflammation, PIN, microbiome-driven inflammatory changes, and stress-response biology in the prostate gland? While the manuscript identifies network hubs and enriched pathways, it often stops short of explaining what these modules biologically represent or how they may influence prostate cancer development, progression, treatment resistance, or immune evasion.

      There are additional publicly available spatial transcriptomic or single-cell datasets which could be used to validate whether the purported immune-cell-specific genes are genuinely enriched in immune populations adjacent to tumour cells. In the drug repurposing analyses, the current study does not explicitly handle prostate cancer subtypes such as HSPC, CRPC, NEPC, or DNPC and co-treatment with androgen receptor antagonism or radiotherapy.

    1. eLife Assessment

      This useful study explores how macrophage cell-cycle state may influence endocytosis, Mycobacterium tuberculosis uptake, and the intracellular stress experienced by bacteria. While the question is interesting and the experimental approach has promise, the evidence for the central claim that endocytic capacity is specifically regulated by cell-cycle stage is incomplete. The main concern is that fluorescence-based sorting and total-fluorescence measurements likely covary with cell size, so the reported phenotypes could reflect biomass accumulation or other cell-cycle-associated changes rather than endocytic capacity as the causal determinant. As a result, whilst the study raises a hypothesis that is of importance, additional controls are required before the proposed mechanism can be considered well supported.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript- "Cell cycle-dependent variation in endocytosis drives phenotypic diversity in M. tuberculosis" by Subhash et al. demonstrates how host cell heterogeneity shapes intracellular pathogen phenotypes. The central and novel finding of this study (G2-phase cells have higher endocytic capacity and harbour more oxidised Mtb) highlights that a host cell cycle (interphase-driven) changes in endocytic capacity regulate bacterial redox states.

      Strengths:

      Overall, the study is well-executed and conceptually rich, establishing a causal link between host cell cycle progression, endocytic heterogeneity, and M. tuberculosis phenotypic diversity.

      The combination of multiple modalities, including live-cell imaging, flow cytometry, scRNA-seq, and redox-sensitive bacterial reporters, supports these findings and substantially strengthens the biological relevance of the work.

      The writing is generally clear, and the figures are well-organised.

      This work will be of interest to readers across cell biology, microbiology, and infection biology

      Weaknesses:

      However, several central claims are only partially supported, the mechanistic depth is limited, and several experimental and analytical concerns need to be addressed.

      Major Comments:

      (1) The authors demonstrate a correlation between the G2 phase and elevated endocytic capacity. However, the mechanistic link (upstream molecular mechanism) between the cell cycle and endocytic upregulation remains largely unaddressed. The authors speculate that membrane biogenesis during volumetric expansion may drive increased endocytosis and note that lipid biosynthesis genes are upregulated in high-endocytic cells. It would substantially strengthen the paper to test this directly, by examining whether inhibition of lipid biosynthesis (e.g., with fatostatin or cerulenin) selectively reduces the G2-associated increase in endocytic capacity. Alternatively, cyclin-CDK axis perturbations (e.g., CDK1 inhibition with RO-3306 to specifically block G2/M entry) could be used to ask whether cells arrested in G2 maintain elevated endocytosis, helping distinguish cell-cycle-position-dependent from cell-cycle-progression-dependent effects.

      (2) The current data show a clear association between high endocytic capacity and more oxidised Mtb, and the authors (consistent with their prior work) hint at lysosomal delivery as the likely mechanism. However, direct evidence for this in the current paper is limited. An experiment examining phagosomal pH or lysosomal fusion (e.g., using a pH-sensitive reporter or lysotracker) specifically in high- and low-endocytic-capacity cells after infection would help confirm this.

      (3) Temporal resolution of Mtb redox dynamics. The plasticity experiment (Figure 6C-D) is elegant and shows that Mtb redox states revert as host cells divide and daughters enter G1. However, the experiment compares day 0 and day 3 post-sorting, which spans multiple cell divisions. While a finer time resolution (spanning 24h) would establish the causal relationship, the authors could discuss the possibility and consequences of multiple cell divisions between day 0 and day 3 used in the present study.

      (4) Relevance of G2 percentages in differentiated macrophages. In Figure 7 and Supplementary Figure S7, only 4.4-5.7% of THP-1-derived macrophages and 5.7% of BMDMs are in G2. While the authors demonstrate statistically significant differences in Mtb redox states between G1 and G2 macrophages, the biological significance of such a small G2 fraction in a non-dividing population deserves discussion specifically with respect to: a) Are these cells re-entering the cycle? b) Is the G2 designation capturing a distinct functional state rather than active cycling? The authors should include additional markers (e.g., phospho-histone H3 for mitotic cells or BrdU incorporation to test for active S-phase) to characterise this population and clarify its identity and origin in differentiated macrophages, thereby meaningfully informing interpretation.

      In conclusion, this is an important mechanism-driven study that highlights an important link in host-driven bacterial phenotypic heterogeneity. The experiments are thorough, the model is well-supported, and the study has implications for infection biology.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors utilize a combination of techniques to show that macrophage endocytic capacity is partially dictated by cell-cycle stage, that Mycobacterium tuberculosis (Mtb) more readily infects macrophages that are in G2/M -phases, and that bacteria that are internalized by macrophages at different stages of the cell-cycle experience different levels of intracellular stress (as reported by the redox state of the bacteria). Furthermore, the authors provide evidence that terminally differentiated macrophages retain memory of the cell-cycle stage that they were in prior to differentiation, at least in the context of endocytic capacity.

      This work provides evidence for the growing idea that fundamental heterogeneity in both host and bacterial organisms can alter the host-pathogen relationship in important ways. However, based on the current data, I am not convinced that the manuscript establishes endocytic capacity as the causal link between macrophage cell-cycle stage and bacterial state. The main issue is that fluorescence-based sorting for cell-cycle stage is likely to covary with cell size. Larger cells, including those later in the cell cycle, may be more likely to fall into the "high" fluorescence gate, while smaller cells may be enriched in the "low" population. Therefore, the observed phenotypes may still be cell-cycle-associated, but the causal determinant could be a correlated feature of cell-cycle progression rather than endocytic capacity itself. This is a significant caveat because nearly all the data, including the live-cell imaging following individual cells, rely on 'total' fluorescence, which will scale strongly with cell size.

      If the authors' conclusion that endocytic capacity is cell-cycle regulated holds true after appropriate controls, this would significantly advance our understanding of the causal interplay between host cell-cycle state, endocytosis, and Mtb physiology. However, an alternative interpretation is that the observed differences in Mtb uptake and bacterial redox state are associated with cell-cycle stage but are not caused directly by differences in endocytic capacity. For example, they could instead reflect other cell-cycle-linked changes in macrophage physiology, such as cell size, intracellular volume, metabolic state, or some other mechanism important for Mtb pathogenesis. If the authors find that their data are best explained by cell-cycle stage independent of endocytic capacity, this would still represent an important advance. However, in that case, the manuscript should clearly distinguish the association with cell-cycle state from the downstream effector mechanisms, which would remain to be determined.

      Strengths:

      The authors utilize various macrophage models for their studies, which is important considering the variability in macrophage behavior, as well as the growing evidence that differences between mouse and human macrophages are relevant for Mtb infection.

      Weaknesses:

      The most important caveat is the covariance between fluorescence-based reporters and cell size. This concern applies to both the sorting experiments, which directly measure total fluorescence, and the time-lapse microscopy experiments, in which the authors show total fluorescence rather than mean, area-normalized fluorescence in Figure 3C. This could be explained by biomass accumulation alone, rather than by a specific cell-cycle-dependent increase in endocytic capacity. Without distinguishing total signal from concentration or activity per unit cell area/volume, it is difficult to conclude that endocytosis itself is regulated by cell-cycle stage rather than simply scaling with cell size.

      Although the authors provide some evidence that the mean GFP intensity, which more closely reflects concentration, differs between the sorted populations in Figure 3B, they do not report statistics for this comparison. Moreover, this control is not carried through the rest of the manuscript, including in key experiments such as Figure 2B. As a result, it remains difficult to determine whether the observed differences between "high" and "low" populations reflect cell-cycle state specifically or instead reflect differences in total reporter fluorescence driven entirely by cell size.

      The evidence for cell-cycle-dependent effects would be more convincing if the authors included additional controls. For example, they could:

      (1) Plot both mean GFP intensity and total GFP intensity in Figure 3B, ideally alongside an unrelated fluorescent reporter that does not vary across the cell cycle. This would help distinguish changes in reporter concentration from changes driven by cell size or total fluorescence.

      (2) Sort cells based on an unrelated fluorescent marker and test whether the same phenotypes - infectivity, dextran uptake, bacterial redox state, etc. - differ between high- and low-fluorescence populations. If these phenotypes are specific to the cell-cycle reporter and not observed with an unrelated marker, this would strengthen the conclusion that the effects are linked to cell-cycle state rather than to fluorescence intensity, cell size, or sorting artifacts.

    1. eLife Assessment

      This important submission from Ambler and colleagues brings new insights into how torpor conditions may confer resilience in cases of cardioprotection. It has novelty, which can be enhanced by additional in vivo support. The study is backed by solid evidence, and represents a unique interoceptive mechanism of interest.

    2. Reviewer #1 (Public review):

      Summary:

      Torpor can be induced by chemogenetic activation of the medial preoptic area. This activation leads to protection from myocardial infarction in an isolated heart preparation despite normalization of the ambient temperature, thus, in principle, uncoupling hypothermia from torpor-induced neuroprotection. Putative pathways of protection are suggested by proteomic studies.

      Strengths:

      (1) Elegant strategy for inducing torpor in rats.

      (2) Appropriate controls for verifying the neuron transducer.

      (3) Cardiac protection is significant and appears independent of hypothermia.

      (4) Interesting omic strategy to begin to find established and novel pathways mediating organ autonomous torpor-induced protection.

      Weaknesses:

      (1) The study would benefit from using inhibitory chemogenetics of the same neurons to demonstrate that this might make cardiac response to ischemia worse.

      (2) Infecting an area of the brain not known to be involved in torpor would be a useful control.

      (3) In vivo cardio protection seems essential as the validation of the strategy requires support that is in the intact animal.

      (4) The assumption that the positive effects of torpor are mediated via a phosphoproteomic change rather than a translational or transcriptional control mechanism is not established.

      (5) A 40 percent reduction in infarct size may work for genetically identical rats with no co-morbidities, but is unlikely to be significant enough to weather the variability that emerges in humans because of these differences and more. The question is not what the mechanism is, but how do we make it more robust? Overall, this is at best a preliminary data set that requires more experiments to deliver on its immense promise.

    3. Reviewer #2 (Public review):

      Summary:

      Elley and colleagues induced a synthetic torpor-like state in rats (a non-hibernating species) by chemogenetically activating neurons in the medial preoptic area of the hypothalamus. They show that this state substantially reduced cardiac infarct size in an ex vivo ischemia-reperfusion model. They further report that protection persisted when ambient temperature was raised to prevent hypothermia, and used exploratory phosphoproteomics to identify candidate cardioprotective signaling pathways.

      Strengths:

      This is the first demonstration that a torpor-like state is cardioprotective in a species that does not naturally enter torpor, which meaningfully advances the potential clinical utility of synthetic torpor. The experimental design is logical, and the controls are generally appropriate. The characterisation of the responsible neuronal population using ISH against QPLOT markers adds mechanistic depth and supports the cross-species conservation argument. The phosphoproteomic analysis, though exploratory, generates plausible and biologically coherent hypotheses grounded in the hibernation literature.

      Weaknesses:

      The primary weakness is that the central conclusion - that hypothermia is not necessary for cardioprotection - exceeds the evidence. The thermoneutral groups were not demonstrably normothermic (36.4 vs 37.05{degree sign}C, p=0.44 with n=6), core temperature telemetry was absent in the majority of control animals contributing to the infarct endpoint, and the decisive test, i.e., a correlation between individual nadir temperature and infarct size, was never performed. Additional weaknesses include the absence of sex-stratified analysis despite known estrogenic contributions to torpor

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Elley and colleagues describes experiments on the effects of synthetic torpor on ex vivo heart ischemia. The key aspect of the study was the use of viral-vector mediated manipulation of the hypothalamic medial preoptic area (MPA) in rats. They used AAV-CaMKIIa-hM3D(Gq). The authors report that chemogenetic activation of the MPA prior to an ex vivo heart ischemia-reperfusion insult induces cardio protection against infarct size that is independent of prior in vivo hypothermia. Phosphoproteomic analysis of cardiac tissue suggested changes in cell survival and death pathways.

      Strengths:

      This study has important strengths. The idea is novel. The experimental design is appropriately rationalized and fascinating. The manuscript is written and presented concisely.

      Weaknesses:

      The study has important weaknesses in the experimental design and validation of the model.

      (1) The study is based on the use of a DREADD-designed viral vector (AAV-CaMKIIa-hM3D(Gq) -mCherry) that is activated by 2 mg/kg IP injection of CNO. The rationale is to putatively activate the MPA. The authors show no evidence for chemogenetic activation of neurons in the MPA. This could be done using a variety of different approaches, even phosphoproteomics.

      (2) The stereotaxic injections are difficult to precisely and locally place, particularly bilaterally. Figure 2F is only a schematic. It would be better to show actual low magnification brain sections (bregma +0.12 to -0.48) from a representative rat to show the placement of the AAV.

      (3) The control rats were injected with AAV-CaMKIIa-EGFP. Why was EGFP used instead of mCherry for the control?

      (4) Ideally, a mutant non-activatable variant of AAV-CaMKIIa-hM3D(Gq) should have been used for a better control.

      (5) The authors should comment on whether there is any neurotoxicity in the MPA associated with the forced AAV expression of hM3D-Gq.

      (6) Is there any inflammatory pathology seen in the MPA with AAV transduction?

      (7) There are no experiments to show that the systemic torpor is specifically associated with the MPA region. Experiments should be done with injections of AAV-CaMKIIa-hM3D(Gq)-mCherry placed in other brain regions, for example, the nearby nucleus accumbens.

      (8) The mapping of the distribution of neurons responsible for synthetic torpor is not mechanistic enough and is not directly to the point. While excitatory and inhibitory markers are examined, a more interesting and deeper approach would have been to use glutamate receptor antagonists to manipulate the torpor response.

      (9) The ischemia and reperfusion aspects of the Lagendorff method need to be clarified. The isolated hearts are already ischemic after their removal from the rat. The reperfusion aspect is caused by reflow of blood to generate oxidative stress, but in the ex vivo model, is there really reperfusion injury?

      (10) The authors show that whole animal oxygen consumption is reduced in the torpor state. The measurement is crude and most likely reflects the inactivity of the animal's skeletal muscle in the torpor state. A more relevant and direct experiment would be to do oxygen consumption (or Seahorse) assays on extracts of the isolated hearts.

      (11) The authors report that the synthetic torpor induces bradycardia. There is no follow-up on this important observation. The MPA-heart connection is not analyzed. (A) Is the link through cardiovascular centers in the brainstem? (B) Is the torpor-induced bradycardia mediated through increased parasympathetic or decreased sympathetic autonomic tone? Pharmacological experiments could also be done.

    1. eLife Assessment

      This study addresses a recent discovery by others that electroconvulsive therapy (ECT) generates seizure activity and spreading depolarization (SD), reflected in large increases in calcium, which can be followed by imaging calcium fluctuations in neurons. This work is useful. However, the evidence to show that SD, rather than seizures, confers the neuroplastic and other therapeutic effects of ECT is incomplete.

    2. Reviewer #1 (Public review):

      The work corroborates the idea, recently suggested by Rosenthal et al. (2025), that spreading depolarization is involved in the mechanisms of electroconvulsive therapy. Using a mouse model of electroconvulsive therapy and various sophisticated approaches to visualize cortical activity, the authors provide an extensive description of traveling calcium waves induced by electroconvulsive stimulation. The study confirms that the calcium events have properties typical of cortical spreading depolarization and seeks to show that the calcium/SD waves mediate therapeutic and neuroplastic effects of electroconvulsive therapy. The authors find that after electroconvulsive stimulation associated with calcium/SD waves, Fos expression increases widely; in the cortex, this increase is localized to the hemisphere affected by calcium waves. They show that some EEG predictors of the beneficial effects of electroconvulsive therapy correlate with the occurrence of calcium/SD waves. Despite the solid methodology and the study's interesting, its conclusions are not fully supported by the data.

      In particular:

      (1)The title of the paper claims that "electroconvulsive stimulation drives cortical spreading depolarization dependent immediate early gene expression". However, immunohistochemical staining shows that Fos expression increases not only in the cortex but also in many subcortical regions, including the hippocampus and amygdala (Figure 5A). Really, conventional electroconvulsive therapy stimulates nearly the entire brain volume and induces generalized seizure activity that can trigger SD not only in the cortex but also in other brain sites. Therefore, regions beyond the cortex can also drive the effects of electroconvulsive therapy. Next, the authors use Fos staining as a marker of neuronal plasticity. However, Fos is also a marker of preceding neuronal activation. As electroconvulsive stimulation, seizures, and SD are associated with high neural activity, it is unclear whether the observed Fos upregulation results from the prior activation or heralds the subsequent plastic changes. Other markers of neuroplasticity (e.g., BDNF) should also be examined.

      (2) Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy. Cortical SD is also tightly coupled with suppression of neuronal activity in affected regions. Although the authors report that postictal suppression is stronger after stimulations with cortical SDs than without SDs, the cortices affected (ipsi) and unaffected (contra) by unilateral cortical calcium/SD events exhibit identical suppression (Figure 6F). The result contradicts established knowledge in the field. If the calcium events are cortical SDs, they should induce EEG suppression only in the affected hemisphere.

      (3) The study states a beneficial role of calcium/SD waves in ECS effects. However, SD alters numerous aspects of brain function, leading to a range of effects that can underlie side effects as well. Assessment of the behavioral effects of stimulation with and without calcium/SD waves can help clarify the issue.

      The results of the work suggest that cortical SD can contribute to electroconvulsive therapy-related mechanisms and help to optimize the stimulation parameters to achieve maximal therapeutic effect.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses the question of mechanisms underlying the therapeutic effects of electroconvulsive therapy (ECT). Clinical efficacy of ECT in major depression (and other disorders) is well established and has often been assumed to be a direct consequence of seizure activity generated by the current application. However, as the authors point out, this explanation is unsatisfactory. A recent study (Rosenthal et al., 2025) provided evidence that ECT generates a wave of cortical spreading depolarization (CSD) in mice, and initial evidence that similar events were generated in patients undergoing ECT. Based on their observations, Rosenthal et al. proposed that CSD, rather than seizure, may engage plasticity mechanisms that contribute to the brain's clinical response to ECT. The current study adds to that prior work by reporting other consequences of CSD, in addition to sustained Ca2+ elevations. The current study also links EEG characteristics immediately following the ECT with the likelihood of generating a CSD, which can help optimize ECT parameters.

      Strengths:

      An important research topic, linking a large set of rodent studies with a limited clinical EEG data set.

      The data acquisition and analyses appear to be of very high quality, and the main results are well illustrated.

      Association between EEG characteristics linked to good clinical outcome matched by mouse EEG data linked to CSD.

      Characterization of multiple consequences of CSD following ECT in the mouse brain.

      Weaknesses:

      The main characterization of CSD propagation comes from GCaMP Ca2+ measurements, as previously reported (Rosenthal et al., 2025). That prior study also provided key electrophysiological evidence of CSD with a DC shift after ECT in mice (supplemental data). Given the prior evidence for ECT-CSD, the additional measures shown in the current manuscript are fully expected. Thus, the 2-photon imaging of Ca2+ elevations following CSD (Figure 4) is consistent with prior 2-photon imaging studies of CSD, and the complex hemodynamic and pH changes are expected to contribute to propagation of EGFP fluorescence changes (Supplemental Figure 5). These data are well presented, but, contrary to the results section here, these results appear confirmatory rather than necessary to build a case that the key event generated by ECT is a CSD.

      The authors state that "our conclusion that CSD is the primary driver of plasticity is based on its role in driving Fos expression" (line 472). Related to the point above, there is already a very well-established literature showing that CSD leads to rapid and robust Fos expression in rodent cortex, so this is fully consistent with prior work. The prior work, CSD-fos work, should be summarized and/or cited more clearly in the manuscript. Showing that Fos increases only in the hemisphere where there is a large CSD-Ca2+ wave is a clear demonstration of this. While Fos increases can certainly be well linked to plasticity in some experimental paradigms, the implication that Fos increases underlie CSD-induced plasticity and possibly therapeutic effects of ECT is not appropriate. Fos increases after CSD are a reliable marker of the very strong neuronal activation that occurs, but Fos increases are not specific for plasticity and can be activated by challenges that do not generate synaptic plasticity. A range of other gene expression changes have been identified with CSD and may contribute to adaptive plasticity; these could be mentioned alongside speculation about Fos. To support the main conclusions of this paper about CSD driving plasticity via Fos, Fos knockout or knockdown studies are needed, as has been used in prior plasticity studies.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript combines widefield calcium imaging, electroencephalography, 2-photon imaging, and immunohistochemistry in mice to re-demonstrate that electroconvulsive stimulation (ECS) induces a seizure followed by cortical spreading depolarization, as previously shown. The putative novel finding - which is not unexpected - is that ECS is also correlated with increased expression of the immediate early gene cFOS, although this has also been shown previously. The authors speculate that CSD drives cFOS expression, which might contribute to the therapeutic effects of ECT; however, experiments performed do not provide causal evidence for this hypothesis. Instead, the authors use expression of cFOS - a nonspecific activity-dependent gene induced in various pathological and non-therapeutic contexts - as a proxy for plasticity and/or therapeutic effect. Hence, overall, the significance of the findings is limited and primarily serves to replicate prior work, with the evidence evaluated as incomplete.

      Strengths:

      The experiments are generally well executed from a technical perspective.

      Main Weaknesses to be addressed in revision:

      (1) The main findings of this paper are replication experiments of prior work, and thus, the novelty and significance of this manuscript are relatively limited.

      - It is already known that the mean frequency of ECT-induced seizures decays between peak and offset in humans (Stuiver et al. Clin Neurophysiol. 2026 Jan:181:2111439. doi: 10.1016/j.clinph.2025.2111439) and mice (Murakami et al. J Pharmacol Sci 2008 Jan;106(1):78-83. 10.1254/jphs.FP0071453), which the authors re-demonstrate in Figure 1.

      - It has already been demonstrated that ECT in mouse models induces lateralized CSD waves in a manner that depends on stimulation parameters and the initial evoked response during stimulation (Rosenthal et al. Nat Comm. 2025 May 18;16(1):4619. doi: 10.1038/s41467-025-59900-1); the authors replicate this in Figures 1, 2, 3, 6.

      - It is already widely established that EEG and calcium signals are highly concordant in mouse brain physiology, as shown in Figure 1. It is already known that CSD propagates from supragranular to granular and infragranular layers (Zakharov et al. Epilepsia. 2019 Dec;60(12):2386-2397. doi: 10.1111/epi.16390) as shown in Figure 4.

      - It is already known that CSD waves induce cFOS expression (e.g., Dell'Orco et al. Front Cell Neurosci. 2023 Dec 14:17:1292661. doi: 10.3389/fncel.2023.1292661; Hermann and Hossman. Neuroscience. 1999 Jan;88(2):599-608. doi: 10.1016/s0306-4522(98)00249-8) as the authors replicate in Figure 5.

      Minimally, the authors should revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field. There is limited innovation in re-demonstrating that these events are seizures and that they involve spreading depolarization.

      (2) The authors frame their hypothesis that CSD could be a potential mediator of the therapeutic effects of ECT, but they do not measure therapeutic effects or directly test this hypothesis. The principal advancement of the paper is showing that ECT-induced CSD triggers hemisphere-specific cFOS expression as a proxy of plasticity. However, it is already known that CSD induces cFOS expression (as noted above). The observation that cFOS expression was induced only by CSD, not by the initial seizure, is likely a byproduct of the greater activity induced by CSD than by seizure. cFOS expression is nonspecific to plasticity or therapeutic effects and can be triggered by many non-therapeutic interventions. The cFOS data thus do not meaningfully measure therapeutic plasticity. The authors also selectively cite references suggesting that EEG metrics such as seizure duration predict positive therapeutic outcomes, but this link is controversial and not well established in the clinical literature.

      Minor Weaknesses:

      (3) For the n=3 mice used for concurrent 2P imaging with microprism implant, these animals also had ChrimsonR co-expression, but there are no optogenetic studies described in this paper, which is confusing. Yet, this co-expression introduces a significant confound, as GCaMP6 emission (525/50nm band in this study) will overlap substantially with the ChrimsonR excitation spectrum. Thus, the fluorescence emission used to image these neurons may be optogenetically activating them at the same time. Please explain.

      (4) Incision of the cortex for implantation of a prism is a significant cortical injury that likely induces CSD instantaneously and may change the propensity for CSD in subsequent recordings. Please comment on this limitation and address how much time elapsed after surgery before imaging.

      (5) Method details are missing or insufficiently described for location, titer, and injection strategy for 2-photon experiments.

      (6) Given the wide range of parameters used for ECS in mice and ECT in humans, the authors should provide tables for what stimulation parameters were used for each recording. These protocols were chosen manually rather than randomly or systematically, which introduces confounding factors into analyses that use parameters as an independent variable.

      (7) While much of the cFOS staining after unilateral CSD shows hemisphere-specific asymmetry, several regions (piriform cortex, amygdala, thalamus) do appear to have bilateral cFOS expression. Please comment on this.

      (8) The discussion states: "If CSD accounts for plasticity effects, triggering a CSD in a non-seizure context may be sufficient to elicit therapeutic effects. This is supported by the clinical success of ultra-brief stimulation treatments that do not cause seizures, such as rTMS with accelerated protocols, which achieves treatment efficacy on par with ECT for major depressive disorder". Are the authors implying that TMS induces CSD? What evidence supports this idea?

      (9) This statement - "Assuming psychosis is the result of thalamocortical coupling that is too weak in frontal areas of the cortex" (lines 583-585) - may be overly speculative.

    1. eLife Assessment

      This important study establishes a robust live-imaging toolkit to characterize excitatory and inhibitory synaptic dynamics during neuronal development, advancing our mechanistic understanding of synaptic homeostasis and neural circuit maturation. The core findings clarify how stable E/I balance is maintained despite persistent synaptic turnover, with broad implications for developmental neurobiology and neurodevelopmental disorders. The methodology and quantitative data are convincing and well validated, and this work represents a significant advance that will be of significant interest to researchers in synaptic biology, cell imaging and neuroscience.

    2. Reviewer #1 (Public review):

      [Editors' note: all three reviewers confirm that all initial concerns have been fully resolved through comprehensive revisions and supplementary analyses.]

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      Comments on revised version:

      The authors have addressed all my questions/comments. No further questions for this manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomato-Gephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HA-Homer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Comments on revised version:

      The authors addressed all my questions and comments. Their edits have made this paper significantly stronger. I believe that this is an important paper for the field.

    4. Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      Comments on revised version:

      The authors have fully addressed my concerns, and this is now a strong manuscript for the synaptic field.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this valuable study, the authors developed long-term imaging tools to simultaneously monitor the temporal and spatial dynamics of excitatory and inhibitory synapses and reported that excitatory and inhibitory synapses need to develop synergistically during synaptogenesis to maintain balance. While the analysis and quantification of the imaging data are incomplete, there is convincing evidence that the developed tools are feasible. If these tools can function stably in vivo, their applications will be much broader.

      We have completely overhauled our analysis and quantification methods and generated custom-made drift correction and tracking pipelines. Also, we have tested these tools ex vivo.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      We thank the Reviewer for their positive feedback and careful review of our study.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      We also thank the Reviewer for their insights and suggestions. We have added discussion of this important point to the Discussion section. Furthermore, we have tested the applicability of our tools ex vivo (new Figures 1, 4, and 6). While using these tools in vivo for live imaging is the eventual goal, we started in a reduced culture system given the relative simplicity. Our current study now provides a framework for future experiments applying these approaches in more complex in vivo systems.

      Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomatoGephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HAHomer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      We thank the Reviewer for their positive assessment of our study.

      Weaknesses:

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Minor weaknesses of the manuscript include:

      (1) The lack of a characterization of endogenous Homer1-positive excitatory synapses using TKIT.

      We attempted to perform live imaging of endogenous Homer1-positive synapses using the TKIT approach by tagging endogenous Homer1 with mClover3 but encountered low signal/noise while live imaging. This prompted us to focus our current study on live imaging endogenous Gephyrin. Future studies using more robust tags (e.g. StayGold, HaloTag) for TKIT tagging of endogenous Homer1 will likely help circumvent this issue.

      (2) Discussion about other approaches to study excitatory and inhibitory synapses using endogenous proteins (e.g., intrabodies - FingR or nanobodies) should be included.

      This important point was also raised by other Reviewers. We have now significantly expanded the Discussion section, including discussion of this point.

      (3) The activity state of a neuron and/or a synapse might alter the dynamic properties (formation, maintenance, and/or elimination). A discussion on whether the overexpression of Homer1 and/or gephyrin might alter synapse/neuron activity would provide greater interpretability of the results. A discussion of the potential limitations and benefits of the reporter and TKIT approaches would be beneficial.

      We agree and have added discussion of these points to the Discussion section.

      (4) A description and interpretation of the computational approach to calculate particle tracking would be helpful. I found that particle tracking figures, while elegant, are difficult to interpret.

      As discussed in more detail below, we have generated drift correction and particle tracking approaches for the revised manuscript. We now elaborate on these new approaches in the paper.

      We thank the Reviewer again for their very helpful input and suggestions.

      Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      We thank the Reviewer for their positive assessment of our study.

      However, a number of aspects remain to be addressed in order for the study to support the claims made by the authors. First, the novelty aspect of the development of the fluorescently tagged synaptic proteins is unclear, since reporters of this nature are in routine use in many labs. Second, the analysis of the acquired images often seems incomplete, with only example images but no quantification shown, or the distinction between spatial and temporal dynamics appearing unclear. Third, given this incomplete analysis, the interpretations of the authors are not always convincingly supported by the data presented. In conclusion, substantial improvements are required to render the main messages of the study clear and compelling.

      We agree and have incorporated all of the Reviewer’s suggestions in the revised manuscript (please see below).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This is an interesting study. This reviewer has the following questions/comments for the authors:

      (1) Please provide evidence that the gRNAs targeting each gene of synaptic protein have no offtarget effects.

      We now include analysis of off-target effects for the TKIT tools (new Figure S6).

      (2) While structural E/I balance is shown, functional electrophysiological validation (e.g., mEPSC/mIPSC ratios) is absent. It is interesting to know whether the balanced functional structural changes translate to functional?

      We thank the Reviewer for this insightful suggestion and now include these recordings in the revised paper (new Figure 8).

      (3) In lines 217-218, please define thresholds for "stable" vs. "dynamic" puncta (e.g., temporal and spatial criteria).

      We more clearly define our categorization parameters (e.g. new Figure 2).

      (4) In Figure 5B: The low co-localization between endogenously tagged Bassoon and antibodystained Bassoon is likely due to the low TKIT efficiency. Quite a few HA-tagged Basson signals are insensitive to Basson-antibody. The authors are suggested to explain those.

      We thank the Reviewer for identifying this and add discussion to the Results section.

      (5) For the data analysis. If each n represents an independent neuronal culture, should the authors are suggested to provide the number of neurons/dendrites analyzed for each independent culture?

      We have added these important details to the manuscript.

      (6) Regarding the title, the author used the term "coordinated dynamics". This reviewer finds it is a bit over-claim because the stable ratios of the number of excitatory synapses and inhibitory synapses are likely an association, not actively "coordinated". I suggest that the authors rephrase this.

      We agree that we cannot argue that excitatory and inhibitory synapses are causally coordinated in our current study. Their levels are likely associated by either association or direct coupling, which we now discuss further in the first paragraph of the Discussion. We have rephrased the title accordingly.

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions that I think will improve the manuscript:

      (1) Please define Syn1/2 on line 129.

      We have defined this in the revised paper.

      (2) For Figures 2B, C, and 4B, C: are the puncta in panel C from the dendrites in panels B? If so, it would be helpful to identify the ROIs selected in panels C.

      We now include this in new Figure 2.

      (3) For the particle tracking figures, while the ability to track all synaptic puncta is very impressive, it is sometimes difficult to clearly track the lifespan of a synaptic puncta from the current figures. I believe that it would be helpful if the authors selected specific examples of synapses formed, maintained, and eliminated.

      We agree and now include more examples.

      (4) I believe that more detail about the computational approach and analysis for the particle tracking (Figs 2E and 4E) would help the interpretability of the figure.

      This important point was also raised by the other Reviewers. We generated custom tools during the revision that significantly expand the capabilities of our tracking approaches and more clearly describe them in the revised manuscript.

      (5) Similar to the rigorous gephyrin TKIT analysis (Fig. 6), did the authors perform a similar analysis for Homer1c TKIT? This might be valuable to confirm that overexpression of the Homer1 reporter does not indirectly alter synapse dynamics.

      We attempted to perform live imaging of mClover3 TKIT-tagged endogenous Homer1 but encountered low signal/noise with live imaging. We now add discussion that optimization of more robust tags (e.g. StayGold, HaloTag) will likely be necessary for live imaging of different target proteins.

      (6) The tools developed by Garbett et al. have the potential to be broadly utilized in the field to provide new insight into the coordination of excitatory and inhibitory synapses. It would thus be helpful for the authors to include a discussion about the strengths and limitations of the reporter and TKIT methods relative to other approaches used to live image synapses (e.g., intrabodies (FingR and nanobodies)).

      We have now significantly expanded the Discussion to include these important points.

      (7) In the discussion, can the authors elaborate on whether it is experimentally feasible to apply their TKIT labeling of gephyrin and Homer1c in the same neuron to assess the endogenous excitatory and inhibitory synapse dynamics from the same neuron?

      We have added discussion of this point and also proof-of-concept data supporting tagging of two postsynaptic targets within the same neuron (new Figure S5D).

      Reviewer #3 (Recommendations for the authors):

      (1) While the new tools described in the current manuscript can undoubtedly be used for the described purposes, the novelty of these tools is unclear to me. Viral vectors expressing fluorescently tagged versions of Homer1, synaptobrevin, and gephyrin are commercially available, e.g., via Addgene, and they are in routine use in many labs. CRISPR-mediated strategies for this purpose have also been previously reported (e.g., Willems et al. 2020, PLOS Biology; Fang et al. 2021, eLife). It is not clear to me how the tools reported here present a significant improvement over existing resources, other than that they use different fluorescent tags. If this aspect is a central part of the current manuscript, it should be expanded on in the discussion, including a direct comparison with available tools to highlight the novel aspects.

      We agree and have significantly expanded the Discussion to include these important points. Also, rather than argue that our tools are superior to pre-existing approaches, we adjust the text to argue that our tools and analytical approaches have been designed and optimized for the purposes we apply them to.

      (2) In addition to generating new tagged constructs, the authors also state that they have developed new imaging and analysis strategies to facilitate long-term assessment of synaptic dynamics. However, in many figures, they present only sample images, with little quantification to allow assessment of the wider relevance of the imaged synapses. For example, in Figures 2C and 4C, they present one example each of, e.g., a stable, nascent, transient, or eliminated synapse. However, they do not provide any quantification on how frequently any of these events occur, or whether they can be reliably quantified at all. These quantifications (i.e., percentage of each event type across a large population of synapses) would be necessary and should be added to demonstrate that this tool can be used for more than single example images.

      We have generated custom-made drift correction and particle tracking approaches for the revised manuscript. Based on the reviewer’s suggestion, we have quantified the relative frequencies of stable, nascent, transient, and eliminated synapses (Fig 2B-G, Fig3A-F, Fig 5A-F, Fig 7B-C). These metrics greatly enhance the biological interpretation of our results. We have also added a supplemental movie with an example image with corresponding categorized tracks for each puncta type (Movie S3)

      (3) The authors do present an automated visual representation of spatial track length across the neuron, e.g., in Figure 2E and 4E, although this is also not quantified. Moreover, the track lengths appear surprisingly short, despite the authors' claims that their analyses 'highlight the dynamic nature of excitatory synapses over these timescales'. It is not clear to me whether these short tracks are more than just jitter, either in the synapses themselves or in the images due to technical limitations. E.g., in panel 2E, I see very few examples in which the track is not simply centered around one point, but actually expands over a distance. Quantification of the distance between start and end points of the tracks would be important to support the claim that these synapses are dynamic in terms of spatial translocation (if that is what the authors meant). Or if the 'dynamic nature' of the synapses referred to temporal dynamics, it is unclear to me how this information can be gained from the represented tracks.

      We thank the reviewer for these excellent points. To accurately access spatial motion, we drift-corrected our images with a custom correction algorithm to eliminate stage or microscope drift as a source of contaminating motion (See Methods, Movie S2), in addition to collecting time-lapse imaging with Nikon perfect focus. We noticed heterogeneity in our cultures such that some areas contained very mobile neurites, while other remained stationary (Fig. S1). We binned movies into either moving or still neurites and assessed spatial metrics as suggested (Fig. S1A). Consistent with our binning, puncta on moving neurites showed larger net displacement (distance between start and end points), but puncta on still neurites also showed ~1 µm net displacement (Fig. S1D). We also quantified puncta speed and found that puncta on moving neurites generally moved faster (Fig. S1C). We appreciate the reviewer’s insight that track length were surprisingly short, and after employing our drift correction and revised tracking methods, we now see substantially longer track lengths (Fig 2E, Fig 3C & F, Fig S2B & C). We additionally see a large fraction of tracks that persist throughout the imaging session (Fig 2E, Fig S2B & C).

      (4) In Figure 3, the authors now quantify track length, but in this case in the unit 'minutes', from which I would interpret that this is now meant to assess the temporal dynamics rather than the spatial dynamics. The lack of a clear distinction between spatial dynamics and temporal dynamics is very confusing to me, since these are entirely independent measures. 'Track length' to me indicates spatial dynamics, and I would expect the units to be a measure of distance. 'Track duration', which the authors also use in some places, but inconsistently as far as I can tell, makes sense to me for the assessment of temporal dynamics, with the units being a measure of time. I would strongly recommend being very clear about this distinction, since the current representation of the data is very difficult to follow and interpret.

      In addition to new spatial metrics, we have clarified in the text when we are referring to spatial dynamics (distance) versus temporal dynamics (time). As suggested, we use duration when referring to time, and speed or distance when referring to spatial metrics.

      (5) The images from the newly generated CRISPR-based tags in Figures 5-7 are striking and very compelling - these will be very useful tools. However, here too, it seems that the interpretation of the data does not really match the results. All quantification indicates that there is very little change in synapse density or other assessed parameters over the time course of the imaging, and yet the authors emphasize the dynamic nature of visualized synapses. More compelling quantification would be needed to support this claim.

      We have quantified spatial and temporal metrics for live neuron culture imaging for all tools developed including CRISPR-based tags (Figure 7).

      (6) The discussion is extremely short and provides almost no integration of the results of the study into the framework of existing knowledge. Instead, it focuses almost exclusively on unanswered questions and future perspectives, which are also important, but not helpful in interpreting the findings from the current study. The latter aspects should be added to provide essential context for the current findings.

      We agree and have added additional discussion of our current findings to help contextualize their significance.

      We thank the Reviewers again for their positive feedback and insightful input, which has undoubtedly strengthened our study.

    1. eLife Assessment

      This important study shows that regions of the human auditory cortex that respond strongly to human voices are also sensitive to vocalizations from closely related primate species. The evidence is convincing and methodologically strong. The work offers significant insight into the evolutionary continuity of voice processing and would be of interest to researchers studying auditory processing and evolutionary neuroscience in general.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance u

      Comments on revised version.

      I thank the authors for having carefully considered and implemented my remarks on the first version.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding-that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations-suggests a potential neural sensitivity to calls form phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weakness:

      The authors only tested vocalizations from three non-human primate species other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      Comments on revised version.

      I have no further comments.

    4. Reviewer #3 (Public review):

      Summary:

      Using fMRI, the authors demonstrate that human temporal voice areas (TVA) respond not only to human vocalizations but also to those of other primates, particularly chimpanzee calls, which share acoustic features with human voices. These findings provide compelling evidence for cross-species vocal processing in the human auditory system and carry important theoretical implications for understanding the evolutionary underpinnings of speech perception.

      Strengths:

      The study offers a valuable comparative design, rigorous acoustic and phylogenetic modeling, and consistent evidence that bilateral anterior TVA regions respond more strongly to chimpanzee vocalizations than to other species' calls. The inclusion of both great apes and monkeys provides a rare cross-species perspective.

      Weaknesses:

      Minor limitations include the acoustic-phylogenetic confound (which the authors partially address with additional analyses), the lack of non-vocal controls to establish true selectivity.

      Overall, the methods, data, and analyses broadly support the claims, with only minor weaknesses that do not undermine the main conclusions. The findings are valuable for the subfield of auditory neuroscience and comparative cognition, with solid evidence supporting the primary claims.

      Comments on revised version.

      After revision, this work has shown great improvement in data analysis, figure organization, and writing. I have no further suggestions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance understanding of the neural bases of voice perception and the evolutionary roots of voice sensitivity in the human brain.

      Weaknesses:

      (1) Acoustic-phylogenetic confound: The design does not fully disentangle acoustic similarity from phylogenetic proximity, as species co-vary along both dimensions. A promising way to address this would be to include an additional model focusing on the acoustic features that specifically differentiate bonobo from chimpanzee calls, which share equal phylogenetic distance to humans.

      (2) Selectivity vs. sensitivity: Without non-vocal control sounds, the study cannot determine whether TVA responses reflect true selectivity for primate vocalizations or general auditory sensitivity.

      (3) Task demands: The use of an active categorization task may engage additional cognitive processes beyond auditory perception; a passive listening condition would help clarify the contribution of attention and task performance.

      (4) Figures and presentation: Some results are partially redundant; keeping only the most representative model figure in the main text and moving others to the Supplementary Material would improve clarity.

      We thank the reviewer for contributing to the improvement of the present study and for the extremely constructive criticism. Concerning the identified weaknesses of our work, we provide here some general answers while the detailed review (below) addresses point-by-point the reviews in high detail.

      (1) We totally agree that acoustics and phylogeny cannot be disentangled in our study, which is a limitation. We now provide the suggested analysis on the acoustic specificities of chimpanzee and bonobo calls.

      (2) This point on selectivity vs. specificity is indeed crucial, and we now provide a more careful viewpoint and phrasing on this aspect, since our study can only provide partial arguments for this important distinction.

      (3) Task demand following species categorization might rightfully yield to the engagement of distinct brain network compared to merely listening to the stimuli. We discuss this aspect and put forward the argument that, while we cannot control for this aspect, our attentional control study performed by an independent sample, N=28 provides clear evidence that no species triggered an attention bias. In other words, task demand might play a role, but at least in the study we know that attentional resources were not biased towards one species in particular since no effects were observed.

      (4) We agree that results were not articulated in a clear fashion and that figures were redundant. We addressed this aspect and regrouped the figures where appropriate while we include the rest in the supplementary material now.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding - that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations - suggests a potential neural sensitivity to calls from phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weaknesses:

      The interpretation of the findings in this paper regarding the evolutionary continuity of voice processing lacks sufficient evidence. A simple explanation is that the observed effects can be attributed to the similarity in low-level acoustic features, rather than effects specific to phylogenetically close species. The authors only tested vocalizations from three non-human primate species, other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      We want to thank the reviewer for the constructive criticism and for evaluating the manuscript.

      Concerning the principal weakness highlighted, we provide new analyses behavioral, acoustics, model-based fMRI that improve our understanding of the influence of both phylogeny and bioacoustics in our data. We argue that the explanation proposed by the reviewer cannot explain our results, as also observed in several other research from us and others. We discuss this aspect and emphasize that including stimuli from more species would greatly improve the understanding of phylogeny and bioacoustics in this context.

      Reviewer #3 (Public review):

      Summary:

      Ceravolo et al. employed functional magnetic resonance imaging (fMRI) to examine how the temporal voice areas (TVA) in the human brain respond to vocalizations from different nonhuman primate species. Their findings reveal that the human TVA is not only responsible for human vocalizations but also exhibits sensitivity to the vocalizations of other primates, particularly chimpanzee vocalizations sharing acoustic similarities with human voices, which offers compelling evidence for cross-species vocal processing in the human auditory system. Overall, the study presents intellectually stimulating hypotheses and demonstrates methodological originality. However, the current findings are not yet solid enough to fully support the proposed claims, and the presentation could be enhanced for clarity and impact.

      Strengths:

      The study presents intellectually stimulating hypotheses and demonstrates methodological originality.

      Weaknesses:

      (1) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy.

      We thank the reviewer for evaluating our manuscript and for the constructive criticism as well as the many suggestions. Concerning the weaknesses of the study, we provide here some quick answers while more detailed responses can be found below.

      (1) We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator.

      (2) We totally agree that figure redundancy was a problem and we now reduced confusion by combining congruent aspects while pushing other results to the supplementary material.

      Recommendations for the authors:

      Reviewing Editor Comments:

      With additional analyses and discussions, the work has the potential to offer important insight into the evolutionary continuity of voice processing.

      We thank the Reviewing Editor for this additional motivation and for offering us the possibility to revise our manuscript. We will now provide our point-by-point reviewing, referring to manuscript modifications by section and/or line number(s). All modifications are also highlighted in light grey in the text.

      Reviewer #1 (Recommendations for the authors):

      The manuscript is clearly written and addresses an important comparative question about the specificity of human TVA responses. The acoustic analyses are well designed, and the imaging work is careful and thorough. However, several conceptual and methodological issues need clarification or tempering of claims, particularly regarding (i) the distinction between sensitivity and selectivity, (ii) the confounding of acoustic and phylogenetic factors, and (iii) the interpretation of "chimpanzee-specific" TVA activity.

      (1) Introduction

      Line 48: cite more recent infant EEG evidence for early voice sensitivity (Calce, Curr Biol).

      The reference and explanation were added, lines 46-48.

      Line 53: mention recent data on voice processing in marmosets (Jafari, Cell Rep; Dureux, Curr Biol).

      We added the references and the mention of these interesting studies on common marmosets, lines 53-54.

      Line 59: Fecteau et al. (2004) already explored cross-species selectivity; please integrate and discuss.

      We now mention here the work from Fecteau and colleagues and its relevance, see lines 57-59.

      Line 70: clarify that in [27] (Bodin et al., 2021) human TVA responded similarly to human nonverbal vocalizations and macaque coos, likely due to acoustic similarity.

      We added this important aspect, thank you for this precision. See lines 71-72.

      Clarify why an active species-categorization task was chosen instead of passive listening, which is standard in TVA research. Were participants familiarized with stimuli beforehand?

      We added a sentence on this aspect, but basically to summarize it here: we wanted to be able to test human recognition of nonhuman primate species’ calls. From the start, we wanted to test the frontal mechanisms related to decision-based processes of humans when categorizing non-human primate calls hence the 2023 article we published. See lines 75-77 and we also added information on familiarization to the stimuli in the Methods, lines 679-682.

      The 16 acoustic features mentioned should be briefly defined earlier, as they are central.

      We feel like describing 16 acoustic parameters in the introduction would be heavy on the reader, so we instead added a reference to the supplementary table (Table S1) in which these are named and described. See line 80.

      Explain why only chimpanzees and bonobos were selected among the great apes, and discuss the value of including both, given their equal phylogenetic proximity but largely dissimilar acoustics.

      The stimuli were obtained by Thibaud Gruber and his team and through collaborations with Katie Slocombe and Zanna Clay. Unfortunately, at the time we could only use chimpanzee and bonobo calls for the great apes. Therefore, it was mainly a material constraint rather than a deliberate choice to exclude other great apes. We now discuss this aspect and present the absence of other great apes as a limitation (lines 587-591).

      Rephrase references to "recruitment" of TVA - this term implies general activation, while the key question concerns selectivity (stronger responses to voices vs. non-vocal controls).

      We rephrased throughout the manuscript, thank you for this suggestion.

      The hypothesis section should more clearly separate the acoustic and phylogenetic predictions, and clarify which earlier data motivate each.

      We now explicitly categorize the hypotheses according to either Bioacoustics or Phylogeny to clarify. We also added references motivating each hypothesis. See lines 114-120.

      (2) Methods

      Clarify whether stimuli were RMS-normalized or otherwise balanced for energy (line 128).

      Sound pressure level was kept constant but the stimuli were not normalized, specifically to avoid a negative impact on their naturality. We added a sentence (lines 131-132) including a reference on this aspect.

      The task design could benefit from reporting accuracy in addition to reaction times for the 4AFC species classification task.

      We agree this aspect was missing. We now report accuracy data (controlled for reaction times and acoustics of Model 3) for the species categorization task (lines 147-165; Fig.1B), and in the Methods (lines 769-786). The fitted regression values of this analysis are also used for a new fMRI model (Model 4), to uncover within-TVA correlates of the probability of correct species categorization (lines 309-325; Fig.4).

      Please note that previously, the behavioral data of the species categorization task were completely absent (N=23), and the reaction times data previously part of Fig.1 were for the species attentional bias task (independent sample of N=28). Since this aspect was not clear at all (same remark by all reviewers—apologies for that), we now include a clear separation in Fig.1, with newly added panels D & E part of a distinct figure area named: “Control task: Testing for Species attentional bias (N=28)”. Panel D illustrates the control task paradigm (each species as exogenous cue; “dot-probe” paradigm) while panel E shows the results (target sine wave tone or “bip” detection reaction times), showing that no species triggered more attentional capture than the others (Species effect non-significant).

      The acoustic parameters used in Models 2 and 3 should be explicitly listed in the Methods (even if already published elsewhere).

      In addition to their description in Table S1, we now include the 16 acoustic parameters used to calculate acoustic distance between the species in the Methods, see lines 828-844.

      Consider simplifying the presentation of the three models: a figure summarizing their relationships would help.

      We now include only one figure (Fig.2) for Model 3, and we pushed model 1&2 to the supplementary material. We also simplified Fig.3 for a clearer view of the overlaps between the 3 models within the TVA.

      The description of “systematic and thorough control of phylogeny” (line 119) is overstated, given that only three nonhuman species were included.

      We agree with the reviewer and we suppressed both “systematic” and “thorough” from the sentence.

      Provide rationale for not including a nonvocal control category (e.g., scrambled vocalizations or environmental sounds) to assess TVA selectivity.

      The main objective of the study was to uncover whether human participants could recognize the vocalizations from nonhuman primates—from both great apes and monkeys—as compared to the human voice. We therefore did not include nonvocal or noise stimuli. We added this point as a limitation in the Discussion (lines 593-596 and 609-611).

      Even though we did not include such stimuli for the reason mentioned above, the delineation of subtypes of nonvocal material within the TVA of our participants (Fig.2) are, in our opinion, clarifying the message: chimpanzee-selective activations are fully within ‘voice vs. animal’ and ‘voice vs. nature’ TVA subareas, while it is not the case in ‘voice vs. music’ and ‘voice vs. noise’ TVA subareas.

      Clarify if participants were trained or had a practice session to recognize the four species before scanning.

      The participants were indeed trained on 3 stimuli per species before entering the MRI scanner. These stimuli were discarded from the species categorization task. We added a sentence about this aspect, see lines 131-132.

      Specify what is meant by "no good or bad response" in the attentional control task (line 724).

      We suppressed this wording as it was highly confusing.

      (3) Results

      Behavioral accuracy should be reported to complement reaction times.

      We now added behavioral data for the species categorization task as well as the neural correlates of accurate species categorization. See our previous response above (‘‘‘).

      Figures 2-4 largely overlap; consider merging or simplifying to reduce redundancy.

      We agree and this point was raised by the other reviewers as well. Task-based results are now presented only for Model 3 as Fig.2, while Fig.3 (previously Fig.5) summarizes the overlap between the three models. Figures for Models 1 & 2, previously labelled Fig.3 and Fig.4, were moved to the supplementary material.

      Figure 2: Please indicate more clearly where "chimp-selective" areas are located (perhaps with zooms).

      We agree, we now modified Fig.2 with zoomed-in panels and a clearer outline of chimp-selective areas (solid blue outline). This outline is also referenced in the text (lines 236-237).

      Correction for multiple contrasts: With many pairwise tests, adjustments (Bonferroni or FDR) should be mentioned explicitly.

      We now specify ‘FDR correction at the voxel level’ at the beginning of the Results section (lines 195-198) as well as in each figure.

      Replace "specific to chimpanzee" with "selective for chimpanzee" to avoid implying exclusivity.

      We made the suggested replacement throughout the manuscript.

      Discuss whether the small macaque-related clusters might simply reflect acoustic overlap rather than true category selectivity.

      We added a section on this important aspect, including results that support the role of mid-STG/STS regions for more noise-like stimuli, including the use of macaque coos. See lines 450-461.

      (4) Discussion

      The discussion overstates claims of "chimpanzee-selectivity" in TVA. The evidence shows relative preference, not absolute selectivity.

      We now specify from the start of the Discussion that we are not interpreting the results as absolute selectivity but rather as more relative preference, see lines 371-373.

      The authors repeatedly conflate acoustic and phylogenetic factors; this should be explicitly acknowledged as a limitation.

      We agree, and we completed the limitations section already dedicated to this aspect by a more explicit account of the confound, see lines 609-611.

      Clarify what is meant by "recruitment" and "selectivity" (lines 411-419, 577). TVA activity often reflects enhanced responses to voices compared to non-vocal sounds, not exclusive activation.

      We clarified this wording in the Discussion (lines 377-378) and replaced another instance by “activated the […]” to make it clearer what we imply, namely enhanced activity triggered by chimpanzee calls within human TVA.

      The lack of non-vocal control conditions should be discussed as a major interpretive limitation.

      We added this point as a limitation in the Discussion (lines 593-596).

      The statement that "chimpanzee-selective activity" arose in humans who have never been exposed to chimp calls (line 450) invites evolutionary speculation but should be more cautiously phrased.

      We agree, and we rephrased by: “[…] with chimpanzee calls triggering responses in the anterior STG/TVA of our human participants […]”. See lines 432-433.

      The comparison to recent macaque data (Giamundo et al., 2024 PNAS) is crucial: these findings of human-voice-selective neurons in macaques directly parallel the present human-chimp result.

      We agree with the reviewer, and we are hopeful to read similar results for other apes/great apes in the future.

      Reviewer #2 (Recommendations for the authors):

      (1) The primate vocalizations used in this study were recorded in diverse social and emotional contexts, which may have contributed to the observed differences in TVA activation. Since the temporal voice areas are known to be sensitive to affective and socially relevant cues, these contextual differences could confound the interpretation of species-specific neural responses. Therefore, I suggest that the authors conduct a post-hoc analysis to quantify and compare the affective valence, arousal levels, and social contexts associated with each stimulus set.

      We agree that the TVA are sensitive to social—or socially relevant—cues, motivating the very thorough work of the expert reserve personnel on-site to accurately categorize the calls according to the very specific context they were produced in. If the reviewer meant presenting these stimuli to non-expert participants and asking them to categorize the context or valence, we think it would make no sense since the ratings would be completely below chance level and therefore uninformative. The newly added behavior—and model-based fmri—data include this crucial point, a factor that we named ‘Context’ in our analyses. In fact, for each species’ 18 stimuli, we control for agonistic and affiliative production context—split evenly, per species. Also, computing an additional posthoc analysis by splitting the stimuli according to Context would result in too few trials to get sensible and reliable fMRI results.

      That being said, our study targets this specific aspect by extracting the acoustic features that characterize our stimulus set the best, across context-species-valence-arousal, which is exactly what we want. Through the three types of modeling we used—from more simplistic to more elaborate the results converge only for one species: chimpanzee calls.

      We think the addition of behavioral data, model-based fMRI data, and the specific analysis on acoustic differences between chimpanzee and bonobo calls strengthens the message and the validity of our findings.

      (2) Although the author mentioned that the behavioral effects triggered by these vocalizations have been reported previously, the behavioral responses of the participants in the current study are also crucial for our understanding of the results. If the MRI data can be combined with the participants' behavioral responses for comprehensive analysis, the conclusions of this study will be more compelling.

      We agree with the reviewer, and we added the behavioral data—controlling for reaction times, production context and acoustics of interest—and we also included a model-based fMRI modeling of the probability of correct species categorization as Model 4, Fig.4. See, respectively: lines 147-165, Fig.1B; Methods, lines 769-786; Neuroimaging results, lines 309-325.

      (3) I am still not convinced that phylogenetic proximity drives the observed neural selectivity. While chimpanzee vocalizations do elicit stronger responses in anterior STG, the claim that this reflects evolutionary relatedness lacks evidence. If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree with the reviewer that generalizing our results in terms of phylogenetic proximity alone is not a viable option. Including many more primate species including other great apes would be necessary, and we mention this crucial aspect in the limitations section. We also insist in the Discussion on the interdependence between phylogeny and acoustics in our data, since: 1) we cannot fully disentangle these factors here, 2) we cannot attribute our results to either one or the other. See lines 387-390, 410-411, 473-477, 587-591.

      If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree, and nobody could disagree: if an auditory object is extremely similar to the human voice in terms of acoustics, it would therefore potentially activate the TVA. This is exactly our message: in the natural ‘auditory world’, the calls from chimpanzees seem to be among the very few animal auditory signals that are sufficiently close, acoustically, to the human voice and therefore trigger TVA activity. They also happen to be the calls from a species which is phylogenetically the closest to humans with minimal differences with other great apes. Our results are in that sense very aligned with work from the laboratory of Pascal Belin, namely on ‘voice patches’ in the primate brain located in the (anterior) TVA, cited in our manuscript.

      We therefore think our interpretation does not exclude that in the near future, similar results within the TVA could be observed for other auditory objects, and if animal, from a species potentially much more distant phylogenetically or from vocal signals of other great apes.

      We added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as scrambled or spectrum shifted per-species stimuli, which would have made the interpretation clearer identical acoustics but alteration/destruction of the species auditory object. See lines 593-596 and 609-611.

      Reviewer #3 (Recommendations for the authors):

      While the manuscript presents intriguing results, several concerns are raised for further consideration, detailed below.

      We thank the reviewer for evaluating the manuscript and for the constructive criticism and suggestions.

      Major concerns:

      (1) This study claims that bilateral anterior superior temporal gyrus (aSTG) in humans can be specifically activated by chimpanzee vocalizations rather than all other primate species after regressing out relevant acoustic parameters using three distinct analyses. I am wondering if a control stimulus (e.g., scrambled chimpanzee vocalizations) were presented, would the activation patterns in these same temporal voice areas (TVA) exhibit significant differences compared to the natural chimpanzee vocalizations?

      We completely agree with the reviewer, and this point was also raised by the other reviewers. We therefore added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as per-species scrambled or spectrum shifted stimuli, which would have made the interpretation clearer—identical acoustics but alteration/destruction of the species auditory object. See lines 609-611.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy. E.g:

      Figure 1C is the same as Figure S1. In addition, Figure 1C lacks a figure legend and descriptive label.

      The scatter plots in Figures 2D, 2H, 3D, 3H, and 4D, 4H are same as those in Figures S2, S3, and S4. However, some of these duplicate plots even have inconsistent axis labels.

      In several panels, the main figures appear to be summaries derived from the supplementary figures. The authors should organize these figures well to eliminate redundancy.

      Please double-check all the figures to make sure of accuracy.

      We agree that the figures were badly organized and were too crowded and redundant. We now suppressed the redundancy between Fig.1 and Fig.S1, and we reduced fMRI results to one figure for statistical Model 3 while the other models are in the supplementary data—we also justify this decision in the text by highlighting that model 3 is the most elaborate and sensitive one. Fig.3 (previously ‘Fig.5’) shows the overlaps between models and was simplified and clarified as well.

      (3) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task. It is possible that processing vocalizations from certain species requires more cognitive effort or induces higher decision uncertainty. Could the observed neural effects be confounded by the decision-making process itself?

      We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator. We now display these results in Fig.4 and we introduce the motivation factor for including a categorization task rather than more traditional passive listening (lines 75-77), as well as limitations, lines 595-596.

      (4) One interesting attempt of this study is to dissociate biologically salient information in animal vocalizations from their low-level acoustic properties. This presents a fundamental conceptual challenge: how to rigorously disentangle a vocalization's species-specific attributes from its inherent acoustic correlates. More precisely, what essential biological information persists in a species' vocal signal after statistically accounting for all quantifiable acoustic features? I recommend that the authors address it in the discussion.

      We thank the reviewer for this very important comment, and for suggesting we discuss it in the manuscript. We completely agree: we cannot fully orthogonalize species and acoustics, and this aspect relates also more broadly to cognitive and affective neuroscience studies involving vocal material. Namely: “What is an auditory object without acoustics?”

      We included a full paragraph on this aspect, see Discussion, lines 570-584.

      (5) If a brain region, such as TVA, is responsive to both acoustic parameters and biological meanings of animal vocalizations, the method used in this study might be inadequate by setting covariates to zero. It is possible that species information is embedded within a specific acoustic pattern. The current modeling approach may not capture such complex information and could potentially introduce bias when estimating the species effect. I recommend that the authors address this issue in the discussion.

      We thank the reviewer for this point once again, we addressed it in the Discussion, lines 581-584, and also in the section dedicated to study limitations, lines 609-613.

      (6) In the discussion, non-human primate vocalizations are "unreadable" to humans. If this is the case, what is the fundamental perceptual difference between these vocalizations and those from the other animal species? An alternative and highly plausible explanation for the findings is the differential familiarity of the participants with the various species, driven by media exposure (e.g., documentaries) or zoo visits and interactions. The authors need to provide a stronger justification for their control stimuli and directly address, either through discussion or additional analysis, how the factor of familiarity might explain their results better than the proposed "evolutionary distance" hypothesis.

      We now discuss this important aspect, see lines 560-569.

      We thought about doing additional analyses on this aspect but we concluded that we did not have any reliable indicators of familiarity for our participants, and additionally they were all recruited for being ‘unfamiliar’ with great apes or old-world monkeys’ vocalized communication.

      Also, frequent mismatches in the media between images of apes and the associated vocal signals (for instance, the depiction of a chimpanzee but with background audio of macaque coos) are not helping this cause.

      Minor:

      (1) No figure legend and result description for Figure 1.

      Figure 1 has a legend, maybe it was cut out during the uploading process, but it is present and verified now.

      (2) In the main text, three statistical models were referenced. Was the data used in each subsequent statistical model derived from the processed data of the preceding model? Please clearly explain this in the main text.

      We now specify this aspect in the Methods and the Results section to clarify that each model is independent from the others (lines 964-966 and 189-191, respectively).

      (3) In Figure 5, the two dashed lines representing Model 1 and Model 2 are confusing for readers.

      We modified the figure (now Fig.3) and simplified it by removing some outlines and clarifying the colors, therefore improving readability.

      (4) Lack of reaction times in the species categorization task.

      We clarified behavioral data, including the results for the species categorization task and for the control, exogenous cueing task, see modified Fig.1 and behavioral results section of the Results.

      (5) Figures 2, 3, 4, 5, Please keep the font size of the figure title consistent.

      Figure title font size were uniformized.

      (6) Line 201, Line 224, and so on, (EFG) → (E, F, G).

      We modified this aspect in every figure legend, including the supplementary material.

    1. eLife Assessment

      Argunşah et al. investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in shaping the responses to single vs multiple whiskers. Based on the observation of a higher density of SST+ interneurons in the septa, the authors investigate the hypothesis that Elfn1-dependent short-term plasticity shapes these responses. This important study is, however, supported by incomplete evidence; factors restricting the strength of evidence are the limited spatial resolution of the multi-unit activity, as well as the lack of a mechanistic explanation. This provocative and intellectually stimulating hypothesis provides a contribution to work on how different cell types shape cortical representation.

    2. Reviewer #1 (Public review):

      Summary:

      Argunşah et al. describe and investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in the whisker-related primary somatosensory cortex (S1). Upon repeated stimulation, the authors report that the response ratio between multi- and single-whisker stimulation increases in layer (L) 4 neurons of the septal domain, while remaining constant in barrel L4 neurons. The authors attribute this divergence to differences in short-term synaptic plasticity, particularly within somatostatin-expressing (SST⁺) interneurons. This interpretation is supported by 1) the increased density of SST+ neurons in L4 of the septa compared to barrel domain, 2) the stronger response of (L2/3) SST+ neurons to repeated multi- vs single-whisker stimulation and 3) the reduced functional difference in single- versus multi-whisker response ratios across barrel and septal domains in Elfn1 KO mice, which lack a synaptic protein that confers characteristic short-term plasticity, notably in SST+ neurons. Consistently, a decoder trained on WT data fails to generalize to Elfn1 KO responses. Finally, the authors report a relative enrichment of S2- and M1-projecting cell densities in L4 of the septal domain compared to the barrel domain, suggesting that septal and barrel circuits may differentially route information about single vs multi-whisker stimulation downstream of S1.

      Strengths:

      This paper describes and aims to study a circuit underlying differential response between barrel columns and septal domains of the primary somatosensory cortex. This work supports the view these two domains contribute distinctly to the processing single versus multi-whisker inputs and highlight the role of SST+ neuron and their short-term plasticity. Together, this study suggests that the barrel cortex multiplexes whisker-derived sensory information across its domains, enabling parallel processing within S1.

      Weaknesses:

      Although the divergence in responses to repeated single- versus multi-whisker stimulation between barrel and septal domains is consistent with a role for SST⁺ neuron short-term plasticity, the evidence presented does not conclusively demonstrate that this mechanism is the critical driver of the difference. The lack of targeted recordings and manipulations limits the strength of this conclusion: SST⁺ neuron activity is not measured in L4, nor is it assessed in a domain-specific manner. The Elfn1 knockout manipulation does not appear to selectively affect either stimulus condition, domain or interneuron subtype. Finally, all experiments were performed under anesthesia, which raises concerns about how well the reported dynamics generalize to awake cortical processing.

    3. Reviewer #3 (Public review):

      Summary:

      This study investigates the functional differences between barrel and septal columns in the mouse somatosensory cortex, focusing on how local inhibitory dynamics (particularly involving SST⁺ interneurons) may mediate temporal integration of multi-whisker (MW) stimuli in septa. Using a combination of in vivo multi-unit recordings, calcium imaging, and anatomical tracing, the authors propose a model in which Elfn1-dependent synaptic facilitation onto SST⁺ interneurons contribute to the distinct sensory responses to MW input in barrels and septa, enabling functional segregation between these domains.

      Strengths:

      The study presents a thought-provoking and useful conceptual model for understanding sensory processing in the somatosensory cortex. While barrel columns have been widely studied, septal regions remain relatively understudied in mice. If septa indeed act as selective integrators of distributed sensory input, this would suggest a novel computational role for cortical microcircuits beyond the classical view focused on barrels. Although still hypothetical, the proposed model in which SST⁺ interneurons contribute to domain-specific sensory responses between barrel and septal domains is intriguing and opens new avenues for investigating inhibitory circuit mechanisms.

      Weaknesses:

      The primary limitation of this study lies in the spatial and cellular specificity of the recording techniques. The physiological data rely predominantly on unsorted multi-unit activity (MUA) recorded with low-channel-count silicon probes. Because MUA aggregates signals from multiple neurons over a radius of approximately 50-100 µm (comparable to or larger than the width of septal domains in mice), it remains difficult to confidently attribute the recorded activity exclusively to septal versus barrel populations. The authors have now addressed this concern more carefully by reframing their interpretation in terms of "septal-enriched" populations and by providing additional threshold-based analyses suggesting that the principal effects are more robust in Layer 4. These additions substantially improve the manuscript and support a more cautious interpretation of the findings. Nevertheless, the proposed Elfn1/SST⁺ mechanism remains supported primarily by indirect evidence. Although the calcium imaging data provide useful support for stimulus-dependent SST⁺ recruitment, these experiments were restricted to L2/3 interneurons and therefore do not directly test the Layer 4 circuit mechanism proposed to underlie the electrophysiological observations. Direct in vivo cell-type-specific recordings and manipulations in Layer 4 would ultimately be required to establish the proposed mechanism more conclusively.

      Comments on revised version.

      I have read the revised manuscript and overall, I think the authors have addressed my major concerns appropriately. I appreciate the substantially moderated interpretation of the findings and the additional analyses clarifying the limitations of the MUA recordings.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Argunşah et al. describe and investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in the whisker-related primary somatosensory cortex (S1). Upon repeated stimulation, the authors report that the response ratio between multi- and single-whisker stimulation increases in layer (L) 4 neurons of the septal domain, while remaining constant in barrel L4 neurons. The authors attribute this divergence to differences in short-term synaptic plasticity, particularly within somatostatin-expressing (SST<sup>+</sup>) interneurons. This interpretation is supported by

      (1) The increased density of SST+ neurons in L4 of the septa compared to barrel domain,

      (2) The stronger response of (L2/3) SST+ neurons to repeated multi- vs single-whisker stimulation and

      (3) the reduced functional difference in single- versus multi-whisker response ratios across barrel and septal domains in Elfn1 KO mice, which lack a synaptic protein that confers characteristic short-term plasticity, notably in SST+ neurons.

      Consistently, a decoder trained on WT data fails to generalize to Elfn1 KO responses. Finally, the authors report a relative enrichment of S2- and M1-projecting cell densities in L4 of the septal domain compared to the barrel domain, suggesting that septal and barrel circuits may differentially route information about single vs multi-whisker stimulation downstream of S1.

      Strengths:

      This paper describes and aims to study a circuit underlying differential response between barrel columns and septal domains of the primary somatosensory cortex. This work supports the view these two domains contribute distinctly to the processing single versus multi-whisker inputs and highlight the role of SST+ neuron and their short-term plasticity. Together, this study suggests that the barrel cortex multiplexes whisker-derived sensory information across its domains, enabling parallel processing within S1.

      Weaknesses:

      Although the divergence in responses to repeated single- versus multi-whisker stimulation between barrel and septal domains is consistent with a role for SST<sup>+</sup> neuron short-term plasticity, the evidence presented does not conclusively demonstrate that this mechanism is the critical driver of the difference. The lack of targeted recordings and manipulations limits the strength of this conclusion: SST<sup>+</sup> neuron activity is not measured in L4, nor is it assessed in a domain-specific manner. The Elfn1 knockout manipulation does not appear to selectively affect either stimulus condition, domain or interneuron subtype. Finally, all experiments were performed under anesthesia, which raises concerns about how well the reported dynamics generalize to awake cortical processing.

      We thank the reviewer for their careful reading of the manuscript and their balanced assessment of both its strengths and limitations. We acknowledge the reviewer’s concerns regarding the lack of direct, layer- and cell-type–specific recordings and manipulations of SST<sup>+</sup> interneurons, as well as the use of anesthesia. As noted in the Discussion, these factors limit the extent to which causal mechanisms can be established and the degree to which the reported dynamics can be generalized to awake cortical processing. For this reason, we intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, developmental, physiological, and genetic evidence, rather than as definitive proof. We believe this conceptual framing appropriately reflects the scope of the current data while highlighting clear directions for future work.

      Reviewer #2 (Public review):

      Summary:

      Argunsah and colleagues demonstrate that SST expressing interneurons are concentrated in the mouse septa and differentially respond to repetitive multi-whisker inputs. Identifying how a specific neuronal phenotype impacts responses is an advance.

      Strengths:

      (1) Careful physiological and imaging studies.

      (2) Novel result showing the role of SST+ neurons in shaping responses.

      (3) Good use of a knockout animal to further the main hypothesis.

      (4) Clear analytical techniques.

      Comments on revisions:

      The authors have effectively responded to my initial critiques - I have no further concerns.

      We thank the reviewer for their positive evaluation of our work and for recognizing the novelty of the findings, the careful physiological and imaging approaches, the use of the Elfn1 knockout model, and the clarity of the analytical framework. We are pleased that the reviewer has no further concerns and appreciates the contribution of this study to understanding the role of SST<sup>+</sup> interneurons in shaping sensory processing in the barrel cortex.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the functional differences between barrel and septal columns in the mouse somatosensory cortex, focusing on how local inhibitory dynamics (particularly involving SST<sup>+</sup> interneurons) may mediate temporal integration of multi- whisker (MW) stimuli in septa. Using a combination of in vivo multi-unit recordings, calcium imaging, and anatomical tracing, the authors propose a model in which Elfn1-dependent synaptic facilitation onto SST<sup>+</sup> interneurons contributes to the distinct sensory responses to MW input in barrels and septa, enabling functional segregation between these domains.

      Strengths:

      The study presents a thought-provoking and useful conceptual model for understanding sensory processing in the somatosensory cortex. While barrel columns have been widely studied, septal regions remain relatively understudied in mice. If septa indeed act as selective integrators of distributed sensory input, this would suggest a novel computational role for cortical microcircuits beyond the classical view focused on barrels. Although still hypothetical, the proposed model in which SST<sup>+</sup> interneurons contribute to domain-specific sensory responses between barrel and septal domains is intriguing and opens new avenues for investigating inhibitory circuit mechanisms.

      Weaknesses:

      The primary limitation of this study lies in the spatial and cellular specificity of the recording techniques. The physiological data rely predominantly on unsorted multi-unit activity (MUA) recorded with lowchannel-count silicon probes. Because MUA aggregates signals from multiple neurons over a radius of approximately 50-100 µm (often wider than the typical septal width in mice), this approach makes it difficult to confidently isolate activity originating strictly from within septal domains. The manuscript would benefit from additional analyses to validate the spatial specificity of these recordings, such as systematically varying spike detection thresholds to test the robustness of domain attribution, as suggested by the reviewer. Furthermore, although the authors now appropriately frame their findings in the Elfn1 knockout mice as indirect evidence, it is worth emphasizing that the study lacks direct in vivo, cell-type-specific recordings and manipulations to more definitively test the proposed mechanism.

      We thank the reviewer for their thorough and constructive evaluation of the manuscript and for highlighting both the conceptual strengths of the study and its technical limitations. We agree that the spatial and cellular specificity of unsorted multi-unit recordings imposes inherent constraints on the interpretation of domain-specific activity, particularly given the narrow width of septal compartments in mice. As now clarified in the manuscript, we do not claim absolute cellular specificity of “septal” recordings but rather interpret them as septal-enriched populations. To directly address this concern, we performed additional threshold-based analysis demonstrating that the key domain-specific effects persist selectively in Layer 4 under stricter spike-detection criteria, supporting a local circuit origin of the critical findings. Further, the more stringent detection criteria (Suppl Fig 3A) collapse the divergence seen in Layer2/3 (Suppl Fig 4C), suggesting that this divergence arises in Layer 4, where SST+ interneuron distributions diverge between barrel and septa.

      We further agree that the Elfn1 knockout results provide indirect, rather than definitive, evidence for causal involvement of SST<sup>+</sup> interneurons and therefore intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, physiological, developmental, and genetic observations. We believe this explicitly moderated interpretation appropriately reflects the scope of the current data while establishing a clear conceptual framework and motivation for future studies employing cell-type-specific recordings and manipulations to directly test the proposed mechanism.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Major comments

      (1) Interpretation of "septal" recordings: The authors claim that the activity recorded from electrodes placed in the septa can be confidently attributed to septal neurons. In my previous review, I raised a major concern that such "septal" recordings likely include spikes from adjacent barrels, given the broad spatial resolution of MUA and the narrowness of the septa in the mouse S1. In fact, the intermediate properties observed in septal recordings from wild-type mice could be explained by a mixture of activity from principal and neighboring barrels-an interpretation that contrasts with the authors' conclusion. Upon reviewing the probe model used (A8x8-Edge-5mm-100-200-177), I noticed a discrepancy between the manufacturer's design and the schematic provided in the manuscript. The electrodes are located near the right edge of the probe rather than the center, suggesting that neurons in adjacent barrels could easily be sampled. In my previous review, I therefore suggested alternative approaches, such as calcium imaging, to more convincingly support the authors' claims. However, the revised manuscript does not include new experiments or additional analyses addressing this issue. Instead, the authors argue that using a high spike detection threshold (SD > 7.5) ensures that recorded activity originates from septal neurons, even though this value does not appear particularly conservative, as it was merely adopted from a previous study without justification in the present context. While I agree that a higher threshold may reduce contamination from distant sources, it does not guarantee that only septal neurons contribute to the signal. By nature, MUA reflects activity from multiple neurons within a radius of at least 50-100 µm. To more rigorously support the claim of spatial specificity, I strongly encourage the authors to reanalyze their existing dataset by systematically varying the spike detection threshold and quantifying how the properties and selectivity of detected units change. If neurons closer to the electrode indeed exhibit distinct domain-specific properties, they should become more prominent as the threshold increases. Such an analysis would strengthen the authors' interpretation and improve the manuscript's impact, even in the absence of new experimental data. Alternatively, the authors could revise their claims to acknowledge that the "septal" electrodes likely record from a population that includes septal neurons as well as neurons located at the periphery of principal and adjacent barrels.

      We agree with the reviewer that, by nature, MUA reflects the activity of multiple neurons within a spatial radius and that recordings obtained from electrodes positioned in the septa may include contributions from neurons located at the periphery of adjacent barrels. This concern is further compounded in superficial layers by probe geometry and orientation: given the narrow width of septa and the lateral spread of processes in upper cortical layers, recordings in L2/3 are inherently more susceptible to spatial mixing than those in layer 4, where columns are more compact and cytoarchitecturally distinct. To directly address these issues, we reanalyzed the same dataset using a more stringent spike detection threshold (SD > 9.5), compared to the originally reported SD > 7.5. Importantly, increasing the threshold selectively reduced or eliminated effects in L2/3, while the key domain-specific differences in L4 responses both the differential MW/SW dynamics in wild-type animals and their attenuation in Elfn1 knockout mice remained robust (the new Supp. Fig. 3. In the manuscript). This threshold-dependent dissociation is consistent with the interpretation that the critical effects reported in L4 arise from neurons spatially closer to the electrode and are less influenced by probe orientation or distant sources, rather than reflecting simple mixing of barrel signals. While this analysis does not claim absolute cellular exclusivity of septal neurons, it provides empirical support that the principal conclusions of the study are robust to stricter spatial sampling criteria and are particularly anchored in L4 circuitry. Accordingly, we now explicitly acknowledge in the manuscript that “septal” recordings likely represent septal-enriched populations rather than purely septal neurons, while emphasizing that the persistence of L4 effects under higher spike-detection thresholds strengthens the conclusion that local L4 inhibitory dynamics underlie the reported functional differences between barrel and septal domains.

      The greater sensitivity of L2/3 results to spike-detection threshold is also expected based on both anatomical considerations and probe geometry. Neurons in L2/3 possess broader horizontal dendritic and axonal arbors and participate in more laterally distributed integration across columns, making population signals in these layers intrinsically less spatially focal. As a result, conservative spike-detection criteria preferentially suppress L2/3 effects, particularly when recordings are obtained with probes optimized for deeper layers. Importantly, our two-photon calcium imaging data while similarly limited to L2/3 demonstrate that SST<sup>+</sup> interneurons show locally measurable and stimulus-specific responses at the single-cell level, providing independent support that L2/3 SST<sup>+</sup> activity is stimulus-modulated rather than artifactual. Taken together, these observations suggest that L2/3 results reflect more distributed and integrative network activity, whereas the L4 effects that persist across thresholds are more directly attributable to local circuit mechanisms. This layer-specific dissociation further supports our interpretation that the central findings of the study are driven by local inhibitory dynamics in L4, with L2/3 activity reflecting downstream integration rather than primary domain-specific computation.

      (2) Interpretation of the Elfn1 KO data: The authors' interpretation that Elfn1-dependent facilitation of SST<sup>+</sup> interneurons underlies the differential sensory responses between barrel and septal domains is conceptually appealing and supported by several converging, albeit indirect, lines of evidence. Specifically, the consistent correspondence among the differential activation of SST<sup>+</sup> neurons upon SWS and MWS, the late development of the barrel-septa differences in the responses to SWS and MWS, and the attenuation of this difference in Elfn1 knockout mice lends plausibility to the proposed model. However, it should be emphasized that the data remain indirect: the study does not include direct recordings of SST<sup>+</sup> neuronal activity from the knockout mice, nor cell-type- specific manipulations to demonstrate causal involvement. The mechanistic explanation therefore represents a hypothesis rather than definitive proof. That said, the authors clearly acknowledge these limitations in the Discussion and appropriately moderate their claims by presenting the SST-Elfn1 mechanism as a working model. Given this careful framing, the current manuscript can be regarded as a valuable conceptual contribution that advances our understanding of how inhibitory dynamics may shape temporal processing in the barrel cortex. Further experiments, as mentioned above, will be essential to test the causal role of this mechanism directly.

      We thank the reviewer for this thoughtful and balanced assessment. We fully agree that the Elfn1 knockout experiments provide indirect rather than definitive evidence for a causal role of SST<sup>+</sup> interneurons in mediating the domain-specific MW/SW response dynamics between barrels and septa For this reason, throughout the revised manuscript we explicitly frame the Elfn1–SST mechanism as a working model rather than a proven mechanism.

      Minor comments:

      The authors have adequately addressed my previous minor comments. In this round, I carefully reviewed the revised manuscript and identified several issues related to references. I would also like to add a brief comment regarding the Discussion section:

      (1) Stachniak et al., 2021 is included in the reference list but is not cited anywhere in the main text. Please either remove this entry or cite it appropriately in the manuscript.

      Removed.

      (2) Yamashita et al., 2018 is cited in the main text (Line 767), but it is not included in the reference list.

      Fixed.

      (3) Sylwestrak and Ghosh, 2012 is cited at Line 261 and Line 270, but likewise absent from the reference list.

      Fixed.

      (4) At Line 497, Chen et al., 2015 is cited, but, the appropriate and original reference would be Chen et al., 2013 (PMID: 23792559), which should either replace or precede the 2015 citation.

      Added.

      (5) At Line 221, El-Boustani et al., 2018 is cited. However, this study is based on the visual cortex, whereas the manuscript concerns the barrel cortex. A more relevant citation (e.g., Lefort et al., 2009 [PMID: 19186171]) would better support the discussion of cellular organization in the barrel cortex. Please consider updating the citation.

      Thank you for this suggestion. We agree with the reviewer and now we have changed El-Boustani with Lefort et al. 2009 as suggested by the reviewer.

      (6) Furthermore, Chakrabarti & Alloway (2006) performed tracer-based mapping of projections from barrel and septal columns in rat S1 and similarly suggested differential organization of M1- and S2projection neurons in the barrel and septal regions.

      Although the current study thoroughly analyzed the layer-specificity of the location of these projection neurons, the lack of explicit discussion of this relevant prior work is a notable omission.

      The authors should incorporate a comparison with these results to better contextualize their findings.

      The following text is added to the discussion: “Our retrograde labeling data supports and expands on previous work proposing similar models (Alloway, 2008; Chakrabarti and Alloway, 2006).”

    1. eLife Assessment

      This potentially valuable study aims to investigate neural correlates of spatial attention in whisker somatosensory cortex (S1) in mice, finding increased sensory-evoked spiking when the mice appear to be attending to the contralateral whiskers. Although some of the results appear to be robust despite relatively small effect sizes, overall the findings are incompletely supported, because attentional modulation is insufficiently distinguished from learning of stimulus-response contingencies, and because the analyses do not adequately consider orofacial movements that may contribute key confounds.

    2. Reviewer #1 (Public review):

      The paper uses a passive whisker detection task in mice to identify a behavioral phenomenon that can reasonably be interpreted as spatial attentional capture. The attentional effect occurs transiently after a successful whisker stimulus detection yields reward, and lasts for a few trials before subsiding. The attentional effect is to the right or left whiskers, depending on whether right or left whiskers are rewarded; no finer spatial resolution for attention was tested. By recording whisker-evoked spiking from single units in S1, the authors show that this form of spatial attention increases the gain of whisker-evoked neuronal responses in S1 for a large subset of S1 units. In contrast, neural responses are not modulated by overall task engagement. Together, these findings show a neural signature of spatial attention in S1 cortex. Because whisker or facial movements were not tracked, it is not clear whether this represents covert attention or whisker movement in response to previously rewarded stimuli, which would be a form of overt attention.

      Substantial attentional modulation of neural responses was observed for a subset of whisker-responsive S1 units, but the effect size was small on average for the total unit population. The top 25% of units showed a ~12% attentional response modulation (relative to firing rate range for each unit), but the median unit showed only a 1.3% response modulation. It would have been useful to analyze the magnitude or prevalence of attentional modulation across layers or in fast-spiking vs. regular spiking units, but this was not reported.

      Major

      (1) It is hard to interpret the underlying causes of the attentional modulation of neural activity without having measured whisker and facial movement. This is a particular issue in S1, where whisker movement against the stimulation grid can alter the mechanical efficiency of stimulus delivery. Such movements would represent overt attention, which would engage an entirely different neural mechanism than covert attention.

      (2) An interesting debate is whether the behavioral phenomenon is best described as attention or as dynamic learning of the stimulus-response association for that block. In Posner-type cued attention tasks, and also in many block-type attention tasks in rodents, animals receive reward for successfully detecting either cued or uncued stimuli, and thus attention (higher response probability or improved psychometric sensitivity for cued stimuli) is at least partially dissociated from the stimulus-reward contingency. That is not the case here. The fact that mice have difficulty learning the contingency reversal suggests that the phenomenon is better explained by attention than by learning the contingency; however, to prove this clearly, the existence of the attentional effect on neural activity in Block 1 vs. Block 2 would have to be shown.

      (3) Some of the graphical representations of the attentional modulation of neural activity are unclear. The single-unit example of attentional modulation is quite strong (Figure 3d). The mean response for the top 25% of units is also visually clear (Figure 3f). But the effect is not apparent at all in Figure 3e, which the figure legend says shows every unit. What is the yellow point and line in this figure? Why isn't the attentional effect visible in this panel? Perhaps I am misunderstanding Figure 3e, but it is not clear to me why it compares Pref>0.5 to Pref<0.5, when the intended analysis suggests it should be Pref>0 to Pref<0? Also in Figure 3, it is critical for the reader to know whether panels 3g-3h represent the top 25% of units or all units. Neither the results text nor the legend is clear on this.

      (4) There is a missed opportunity to quantify attentional modulation across cortical layers, since laminar probes and Neuropixels probes were used for the recordings. In addition, there is no separation of fast-spiking from regular-spiking units, and no quantitative metrics are provided to assess the quality of single units. This could reveal key aspects of cortical processing of attentional signals.

    3. Reviewer #2 (Public review):

      Summary:

      Dyce et al investigate the modulation of sensory responses in the somatosensory 'barrel' cortex during a novel whisker vibration detection task in head-fixed mice, aiming to find correlates of spatial attention in both the animals' behavior and their neuronal activity.

      Strengths:

      The authors produced an extensive and parameterized dataset of both behavioral responses and neuronal activity, with >3000 single units of which >1400 were responsive.

      Weaknesses:

      In my view, the main conclusions of the manuscript are not currently well supported by the data.

      The authors effectively define "spatial attention" as a state where an animal responds more to a stimulus that gives more rewards (out of two possible stimuli presented on different sides of the snout, i.e., segregated spatially). If one defines spatial attention purely in these terms, then their findings do show neuronal correlates of spatial attention. However, those neuronal correlates can be explained by known aspects of neuronal responses in the barrel cortex.

      This plays out in several different ways:

      From the behavioral point of view, greater attention may correlate with an increased hit rate to stimuli on the rewarded side, but in the absence of other supporting measurements, the relationship could well be the opposite: an animal could pay more (rather than less) attention to the stimulus delivered on the unrewarded side, to make sure it suppresses the incorrect response. It is impossible to tell, as the data don't provide an independent measurement of whether the animal is paying greater attention to, or is more aware of, one side than the other, nor do they provide an independent measurement of neuronal tuning on either side. There is no separate measurement of arousal either (e.g., via pupillometry or locomotion).

      The experimental design involved two blocks on each daily task session, with the second block reversing the side on which rewarded stimuli were delivered. Reinforcing one's doubts about the behavior and its interpretation, mice had much poorer performance on each day's second block, to the extent that perceptual sensitivity (d') was the same for both sides: d' did not increase after reward reversal for stimuli on the initially unrewarded side. This further emphasizes the lack of a separate demonstration of focused "spatial attention".

      Much of the data (both behavioral and neuronal) could be accounted for, e.g., by a strategy where the mouse keeps a token in working memory of what side seems to be driving rewards, while maintaining equally strong sensory drive on both sides, but with no attentional shift at all. The policy would be to respond more whenever the stimulated side matches the token in memory (thus also reinforcing the token, thus enhancing performance next time). This would be easily implemented with a disinhibitory reward-modulation signal such as the one multiple researchers have found carried by VIP neurons (e.g., Szadai et al DOI: 10.7554/eLife.78815).

      Similarly, the fact that "attended trials" (Pref > 0) produced greater responses than "unattended trials" appears to be explainable as follows. Here, "attended" trials are those where the contralateral stimulus is presented (and, if responded to, is rewarded), "unattended" trials are those where the stimulus is ipsilateral (and not rewarded). The animal responds more (at least in the first block) to stimuli delivered to the contralateral pad - i.e., rewarded as opposed to unrewarded ones. Beyond the knowledge mentioned above that cortex-wide VIP sensitivity to rewards can drive disinhibition in general, activity modulation dependent on rewards and outcomes (and stimulus value) has been established specifically in the barrel cortex (e.g., Lacefield et al DOI: 10.1016/j.celrep.2019.01.093, Bale et al DOI: 10.1016/j.cub.2020.10.059, Banerjee et al DOI: 10.1038/s41586-020-2704-z, Chereau et al 10.1038/s41467-020-17005-x). The reward- and value-evoked activity demonstrated in those papers would suffice to predict more activity at the contralateral electrode on "attended" trials, along the lines of the findings in Ramamurthy et al (DOI: 10.1038/s41467-025-60592-w) and consistent also with the enhanced "attentional modulation" on hit trials.

      Other aspects of the analysis and terminology lead to confusing outcomes. For example, in the analysis in Figure 3, Performance averaged in a set of trials around a given trial is defined as the mean rate of responses to stimulation on either side - regardless of whether those responses are correct (since the stimuli can be on either side, but only one side is correct and gets rewarded and putatively reinforced). Thus, this definition of "Performance" can increase with the rate of incorrect licks to the wrong side and is at odds with the normal use of the word. On trials where this Perf = 1 and the stimuli are balanced on either side, this corresponds to a true performance (and reward rate) of only 0.5 - what one would normally consider random discrimination between the sides. Thus, Perf = 1 trials may still give a low reward rate and, if responses scale with reward, a small effect of reward. Hence, based on known properties of reward dependence, greater correlation of neuronal activity with "Preference" than with "Performance" would be expected, rather than reflecting a new aspect of "spatial attention". A definition of performance more in line with established practice and measuring side-to-side discrimination (corresponding more closely to the authors' "Preference" parameter) would have shown this more clearly.

    4. Author response:

      (1) Introduction & Roadmap

      We are grateful to the Reviewers for engaging with outstanding questions relating to our findings’ connections to multiple subdisciplines of cognitive neuroscience. Noting that Reviewers 1 and 2 interpreted our findings differently, we welcome the opportunity to engage in what Reviewer 1 characterised as “an interesting debate”. To promote a shared understanding and discussion of our findings, we have organised our response to address more technical comments first.

      Our provisional response is organised as follows: Section 2 addresses selected technical comments relating to our Results. Section 3 addresses comments related to the design of our behavioural paradigm. Section 4 focuses on the broader interpretation of our findings. Section 5 concludes our provisional response with potential future directions and a summary of the significance of our findings.

      (2) Selected technical comments related to our Results

      We apologise to Reviewer 2 for the confusion in relation to the meaning of “attended” and “unattended” trials. What we said was “Positive Pref values indicate a higher response rate to the contralateral side than the ipsilateral side (relative to the electrode)” (Figure 3c caption), “we indexed all contralateral whisker vibrations according to their associated Perf and Pref” (Results text), and “we divided trials into (contralaterally) attended (Pref<sub>C/L</sub>: Pref>0) and unattended (Pref<sub>I/L</sub>: Pref<0) groups” (Results text). We can confirm that we defined an “unattended trial” (Pref<0) as a contralateral stimulus trial in the centre of an epoch (10-15 trials) within which the mouse responded (licked) more frequently to ipsilateral stimuli. Critically, we did not define an unattended trial as an ipsilateral stimulus trial. Furthermore, attention thus defined (i.e. Pref>0) can vary independently of the whisker stimulus associated with rewards. Indeed, while we initially did not include this result in our paper for the sake of brevity, even unrewarded “attended” trials (Pref>0) evoked significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials. We note that this is an analysis suggested by Reviewer 1, and we will include and discuss this result in our revised manuscript (e.g. in relation to literature suggested by Reviewer 2). For additional clarity, we use “Performance” (Perf) in relation to overall stimulus detection, consistent with the analysis of Lee et al. (2020), which found this measure was correlated with pupil diameter in a vibrissal target detection task.

      We thank Reviewer 1 for noticing that the axes on Figure 3e should be labelled “Pref>0” (Y axis) and “Pref<0” (X axis), as suggested by the figure caption. We will correct this in our revised submission. The yellow point on Fig 3e shows the unit from Fig 3d, while the yellow line in Fig 3e shows the magnitude of that unit’s (non-normalised) gain modulation. While this is alluded to in the Results text (“The example unit in Figure 3d is in the 93rd percentile of units for raw modulation depth (ΔHits(attended – unattended) = 3.3 spikes/second; yellow line in Fig.3e)”, this should be explained in the Figure caption, and it will be in our revised manuscript. We would also like to clarify that Figures 3g–3h display results for all units, not just the top 25%. We agree this is not sufficiently clear and we will rectify this in our revised manuscript. Addressing Reviewer 2, while we acknowledge that mice responded less to both stimuli in the second block, they also meaningfully adjusted their behaviour to the reversal in reward contingencies: their responses to the previously rewarded stimulus reduced significantly more than those to the previously unrewarded stimulus.

      (3) Design of the behavioural paradigm

      We made a deliberate design choice to maximise the ecological validity of our behavioural paradigm, and note that there are advantages to doing so. For example, our paradigm can be used to show that even unrewarded “attended” trials (Pref>0) evoke significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials (see Section 2, above). Indeed, it is precisely this finding that makes our paradigm uniquely suited to the investigation of value-driven attentional capture (Anderson et al., 2011): in this instance attention directed to stimuli that are no longer rewarded despite equal availability of rewarded stimuli. This finding also demonstrates that our paradigm dissociates attention from stimulus-reward contingency at least as well as other paradigms which have been successfully used to study spatial attention in mice. As noted in Section 2, we will discuss this result in relation to other relevant research (e.g. Ramamurthy et al., 2025) in our revised manuscript.

      Briefly, the direct manipulation of reward contingencies is one of two noteworthy methodological distinctions between our own paradigm and that of Ramamurthy and colleagues (2025). The task of Ramamurthy et al. (2025) associated all whisker stimuli with rewards and delivered stimuli to different whiskers on a single whisker pad. These methodological distinctions may have reduced the relevance of the spatial differences between stimuli to the mice undertaking the task. Indeed, it is not certain that a mouse would treat the unilateral variation in whisker stimulation Ramamurthy and colleagues delivered as primarily spatial or featural. The psychophysical and neural differences between spatial and featural attention in humans suggest dissociable underlying mechanisms, and the same may be true in mice. Thus, our own paradigm may more effectively isolate spatial attention from featural attention. Conversely, to the extent that the findings of Ramamurthy and colleagues do reflect spatial attention, our combined findings and paradigms help elucidate the associated mechanisms across spatial scales in mice.

      We acknowledge that spatial cueing is well-suited to isolating the effects of covert attention from other forms of attention. However, it should be noted that spatial cueing in rodents is subject to its own challenges, including limitations in trial numbers due to the required manipulation of stimulus intensity (Reynolds et al., 2000; Herrmann et al., 2010), cue validity and associated trial probabilities (Peterson & Gibson, 2011; Girardi et al., 2013). Such experiments are further complicated by the duration and efficacy of training (i.e. the number of mice that learn the task; Wang & Krauzlis, 2018; Hu & Dan, 2022). It is also worth noting that trial probability manipulations introduce the same limitation in trial numbers with block-type attention tasks (You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022).

      While there are clear differences between our own paradigm and those mentioned above, there are also important similarities. First, these tasks are all goal-directed, stimulus-driven, and reliant on learned task contingencies (e.g. Peterson & Gibson, 2011; Girardi et al., 2013). Furthermore, these paradigms are all operant conditioning protocols which leverage learned stimulus-reward contingencies to train attention-related behaviours in mice. A noteworthy similarity between our findings and those of authors using block-type attention tasks in particular (e.g. You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022) is the observation of apparent attentional biases in behavioural responses independent of the experimental manipulations (i.e. stimulus probability / reward contingency).

      (4) Comments relating to the broader interpretation and discussion of our findings

      Fundamentally, attention involves dedicating limited processing resources to some stimulus events at the expense of others. The design of our behavioural paradigm was informed by existing literature on spatial attention in humans, non-human primates, and mice. Our choice of behavioural and neuronal measures as proxies for attention in mice is consistent with this literature. It is technically possible “an animal could pay ‘more’ (rather than less) attention to the stimulus delivered on the unrewarded side, to make sure it suppresses the incorrect response”, but this seems unlikely given what is known about how attention is typically allocated in such tasks, based on the previously mentioned literature.

      With respect to the interpretation and discussion of our findings, Reviewer 1 describes them as “a behavioral phenomenon that can reasonably be interpreted as spatial attentional capture” but suggests they do not clearly distinguish whether this attentional capture is covert or overt. We respectfully disagree for three reasons. First, as discussed in our paper, whisker motion during detection tasks has consistently been associated with reduced detection performance (Ollerenshaw et al., 2012; Kyriakatos et al., 2017; Vandevelde et al., 2023), suggesting that a “receptive” strategy (Diamond & Arabzadeh, 2013) of whisker immobilisation is more applicable to the current data than a “generative” strategy of asymmetric whisker movement (O'Connor et al., 2010; Dominiak et al., 2019). Second, if our behavioural and neuronal findings were due to the mice moving their whiskers to maximise contact with the meshes, we would expect increased evoked neuronal responses to be associated with greater Perf, not just with greater Pref. This pattern was not observed. Of course, the mice might have employed different whisker movement strategies during epochs of high Pref and Perf, but this seems unlikely and is not a parsimonious explanation for our findings. Third, as noted in the Methods section of the paper, we deliberately positioned the meshes close to the base of the whiskers, limiting the impact of whisker movements on stimulus detectability and the incentive to make them.

      In contrast, Reviewer 2 questions the interpretation of our findings as evidence of spatial attention and suggests they might reflect working memory instead. Current research suggests attention and working memory are intimately related integrative brain functions. Indeed, some researchers have even proposed that working memory might be a form of internally directed attention (Awh & Jonides, 2001; Chun, 2011; Gazzaley & Nobre, 2012; Kiyonaga & Egner, 2013; or vice versa: Libedinsky & Fernandez, 2019). Consistent with the comments of Reviewer 2, more recent work seems to emphasise the coordination of attention and working memory (e.g. Joe & Kim, 2023; Zhu et al., 2026; for reviews see Huynh Cong & Kerzel, 2021; van Ede & Nobre, 2023), along with shared mechanisms (Kiyonaga et al., 2021; Panichello & Buschman, 2021), and nuanced dissociations (Liu et al., 2025). Attention is difficult to dissociate from working memory partly because there are multiple definitions (and/or types) of attention. We did not discuss the various definitions and/or forms of attention at length in our paper, but we will briefly discuss this in the revised manuscript.

      The “interesting debate” to which Reviewer 1 refers could also be described as vigorous, despite approximately three decades of research. This debate broadly relates to the degree to which attentional control is driven by exogenous (e.g. colour contrast) versus endogenous factors (e.g. the focus of spatial attention, see Fig.2 in Belopolsky et al., 2007; see also: Liesefeld & Mueller, 2020; Manini et al., 2021; Beffara et al., 2022), and the degree to which this is a function of experimental context. The review article by Luck et al. (2021) entitled “Progress toward resolving the attentional capture debate” provides a striking illustration of this debate, as do the twenty-two commentaries (and three commentary responses) associated with it. Admittedly, this debate largely revolves around human attention experiments, and human cognition may be more complex than mouse cognition. However, the complexity of human cognition may also be easier to study and appreciate because complex behavioural experiments can be explained to, understood, and performed by human participants with relative ease.

      (5) Comments relating to future directions and the significance of our findings

      The complexity of the attentional capture debate underscores the importance of developing accessible and scalable animal experiments which can be used to provide mechanistic insights. If the human attention literature is any indication, a diversity of rodent experimental paradigms will be necessary to thoroughly map the neuronal implementation of spatial attention. Returning to our paradigm, Reviewer 1 noted that valuable insights into the mechanisms of vibrissal spatial attention might be obtained from comparing the magnitude of attentional modulation we observed between putative regular and fast-spiking categories of units, and between units located in different cortical layers. We agree it is important to understand spatial attention with cell-type and circuit (including laminar) specificity. However, because we could not persuasively cluster our units based on waveform width, and because of the lack of histological data, segregating units on the basis of such variables is not feasible. Despite our assertion that our findings reflect the effects of covert attention (contra Reviewer 1), we agree that future experiments will be required to conclusively rule out overt attention. Noting the proximity of the meshes to the base of the whiskers in our paradigm, and the difficulty of tracking whiskers in this context, Botulinum toxin injections (as in Ramamurthy et al., 2025) might be a means of achieving this.

      The above notwithstanding, our findings provide multiple contributions to the literature on spatial attention (and perhaps working memory). We detected significant attentional gain modulation across a population of 1461 responsive units. While the gain modulation exhibited by the median unit was modest (albeit statistically significant), the top 25% of responsive units showed a ~12% response modulation (relative to firing rate range for each unit), and ~21% of responsive units were suppressed by the average vibrissal stimulus in the unattended state. Our experimental framework offers an accessible platform for future studies leveraging genetic and circuit-level interventions to dissect the cell-type specific mechanisms of spatial attention. Our work is timely, noting the recent focus of human research on the nexus of attention, selection history, and valence (e.g. Serences, 2008; Della Libera & Chelazzi, 2009; Della Libera et al., 2011; van den Berg et al., 2014; Kim & Anderson, 2019, 2023). Our work is also uniquely poised to stimulate new interdisciplinary research into the circuit mechanisms of value-driven attentional capture, with translational relevance to psychopathologies such as ADHD, addiction, and depression; where value-driven attentional capture is altered (for a review see Anderson, 2021).

      References

      Anderson, B. A. (2021). Relating value-driven attention to psychopathology. Curr Opin Psychol, 39, 48-54. https://doi.org/10.1016/j.copsyc.2020.07.010

      Anderson, B. A., Laurent, P. A., & Yantis, S. (2011). Value-driven attentional capture. Proceedings of the National Academy of Sciences of the United States of America, 108(25), 10367-10371. https://doi.org/10.1073/pnas.1104047108

      Awh, E., & Jonides, J. (2001). Overlapping mechanisms of attention and spatial working memory. Trends Cogn Sci, 5(3), 119-126. https://doi.org/10.1016/s1364-6613(00)01593-x

      Beffara, B., Hadj-Bouziane, F., Ben Hamed, S., Boehler, C. N., Chelazzi, L., Santandrea, E., & Macaluso, E. (2022). Dynamic causal interactions between occipital and parietal cortex explain how endogenous spatial attention and stimulus-driven salience jointly shape the distribution of processing priorities in 2D visual space. Neuroimage, 255. https://doi.org/10.1016/j.neuroimage.2022.119206

      Belopolsky, A. V., Zwaan, L., Theeuwes, J., & Kramer, A. F. (2007). The size of an attentional window modulates attentional capture by color singletons. Psychonomic Bulletin & Review, 14(5), 934-938. https://doi.org/10.3758/Bf03194124

      Chun, M. M. (2011). Visual working memory as visual attention sustained internally over time. Neuropsychologia, 49(6), 1407-1409. https://doi.org/10.1016/j.neuropsychologia.2011.01.029

      Della Libera, C., & Chelazzi, L. (2009). Learning to Attend and to Ignore Is a Matter of Gains and Losses. Psychological Science, 20(6), 778-784. https://doi.org/10.1111/j.1467-9280.2009.02360.x

      Della Libera, C., Perlato, A., & Chelazzi, L. (2011). Dissociable Effects of Reward on Attentional Learning: From Passive Associations to Active Monitoring. PLoS One, 6(4). https://doi.org/10.1371/journal.pone.0019460

      Diamond, M. E., & Arabzadeh, E. (2013). Whisker sensory system - from receptor to decision. Prog Neurobiol, 103, 28-40. https://doi.org/10.1016/j.pneurobio.2012.05.013

      Dominiak, S. E., Nashaat, M. A., Sehara, K., Oraby, H., Larkum, M. E., & Sachdev, R. N. S. (2019). Whisking Asymmetry Signals Motor Preparation and the Behavioral State of Mice. J Neurosci, 39(49), 9818-9830. https://doi.org/10.1523/JNEUROSCI.1809-19.2019

      Gazzaley, A., & Nobre, A. C. (2012). Top-down modulation: bridging selective attention and working memory. Trends Cogn Sci, 16(2), 129-135. https://doi.org/10.1016/j.tics.2011.11.014

      Girardi, G., Antonucci, G., & Nico, D. (2013). Cueing spatial attention through timing and probability. Cortex, 49(1), 211-221. https://doi.org/10.1016/j.cortex.2011.08.010

      Herrmann, K., Montaser-Kouhsari, L., Carrasco, M., & Heeger, D. J. (2010). When size matters: attention affects performance by contrast or response gain. Nat Neurosci, 13(12), 1554-1559. https://doi.org/10.1038/nn.2669

      Hu, F., & Dan, Y. (2022). An inferior-superior colliculus circuit controls auditory cue-directed visual spatial attention. Neuron, 110(1), 109-119 e103. https://doi.org/10.1016/j.neuron.2021.10.004

      Huynh Cong, S., & Kerzel, D. (2021). Allocation of resources in working memory: Theoretical and empirical implications for visual search. Psychon Bull Rev, 28(4), 1093-1111. https://doi.org/10.3758/s13423-021-01881-5

      Joe, J., & Kim, M. S. (2023). Spatial Attention in Visual Working Memory Strengthens Feature-Location Binding. Vision (Basel), 7(4). https://doi.org/10.3390/vision7040079

      Kanamori, T., & Mrsic-Flogel, T. D. (2022). Independent response modulation of visual cortical neurons by attentional and behavioral states. Neuron, 110(23), 3907-3918 e3906. https://doi.org/10.1016/j.neuron.2022.08.028

      Kim, H., & Anderson, B. A. (2019). Dissociable neural mechanisms underlie value-driven and selection-driven attentional capture. Brain Research, 1708, 109-115. https://doi.org/10.1016/j.brainres.2018.11.026

      Kim, H., & Anderson, B. A. (2023). Primary Rewards and Aversive Outcomes Have Comparable Effects on Attentional Bias. Behavioral Neuroscience, 137(2), 89-94. https://doi.org/10.1037/bne0000543

      Kiyonaga, A., & Egner, T. (2013). Working memory as internal attention: toward an integrative account of internal and external selection processes. Psychon Bull Rev, 20(2), 228-242. https://doi.org/10.3758/s13423-012-0359-y

      Kiyonaga, A., Powers, J. P., Chiu, Y. C., & Egner, T. (2021). Hemisphere-specific Parietal Contributions to the Interplay between Working Memory and Attention. J Cogn Neurosci, 33(8), 1428-1441. https://doi.org/10.1162/jocn_a_01740

      Kyriakatos, A., Sadashivaiah, V., Zhang, Y., Motta, A., Auffret, M., & Petersen, C. C. (2017). Voltage-sensitive dye imaging of mouse neocortex during a whisker detection task. Neurophotonics, 4(3), 031204. https://doi.org/10.1117/1.NPh.4.3.031204

      Lee, C. C. Y., Kheradpezhouh, E., Diamond, M. E., & Arabzadeh, E. (2020). State-Dependent Changes in Perception and Coding in the Mouse Somatosensory Cortex. Cell Rep, 32(13), 108197. https://doi.org/10.1016/j.celrep.2020.108197

      Libedinsky, C. D., & Fernandez, P. F. (2019). Graded Memory: A Cognitive Category to Replace Spatial Sustained Attention and Working Memory
 Yale J Biol Med, 92(1), 121-125. https://www.ncbi.nlm.nih.gov/pubmed/30923479

      Liesefeld, H. R., & Mueller, H. J. (2020). A theoretical attempt to revive the serial/parallel-search dichotomy. Attention Perception & Psychophysics, 82(1), 228-245. https://doi.org/10.3758/s13414-019-01819-z

      Liu, Y., Fu, Y., Tang, E., Wu, H., Han, J., Xie, M., Zhang, Y., Peng, B., Huang, J., Liu, H., Chen, H., & Qin, P. (2025). Neural dissociation of attention and working memory through inhibitory control. Nat Commun, 17(1), 22. https://doi.org/10.1038/s41467-025-66553-7

      Luck, S. J., Gaspelin, N., Folk, C. L., Remington, R. W., & Theeuwes, J. (2021). Progress toward resolving the attentional capture debate. Visual Cognition, 29(1), 1-21. https://doi.org/10.1080/13506285.2020.1848949

      Manini, G., Botta, F., Martin-Arevalo, E., Ferrari, V., & Lupianez, J. (2021). Attentional Capture From Inside vs. Outside the Attentional Focus. Frontiers in Psychology, 12. https://doi.org/10.3389/fpsyg.2021.758747

      O'Connor, D. H., Clack, N. G., Huber, D., Komiyama, T., Myers, E. W., & Svoboda, K. (2010). Vibrissa-based object localization in head-fixed mice. J Neurosci, 30(5), 1947-1967. https://doi.org/10.1523/JNEUROSCI.3762-09.2010

      Ollerenshaw, D. R., Bari, B. A., Millard, D. C., Orr, L. E., Wang, Q., & Stanley, G. B. (2012). Detection of tactile inputs in the rat vibrissa pathway. J Neurophysiol, 108(2), 479-490. https://doi.org/10.1152/jn.00004.2012

      Panichello, M. F., & Buschman, T. J. (2021). Shared mechanisms underlie the control of working memory and attention. Nature, 592(7855), 601-605. https://doi.org/10.1038/s41586-021-03390-w

      Peterson, S. A., & Gibson, T. N. (2011). Implicit attentional orienting in a target detection task with central cues. Conscious Cogn, 20(4), 1532-1547. https://doi.org/10.1016/j.concog.2011.07.004

      Ramamurthy, D. L., Rodriguez, L., Cen, C., Li, S., Chen, A., & Feldman, D. E. (2025). Reward history guides focal attention in whisker somatosensory cortex. Nat Commun, 16(1), 5580. https://doi.org/10.1038/s41467-025-60592-w

      Reynolds, J. H., Pasternak, T., & Desimone, R. (2000). Attention increases sensitivity of V4 neurons. Neuron, 26(3), 703-714. https://doi.org/10.1016/s0896-6273(00)81206-4

      Serences, J. T. (2008). Value-Based Modulations in Human Visual Cortex. Neuron, 60(6), 1169-1181. https://doi.org/10.1016/j.neuron.2008.10.051

      van den Berg, B., Krebs, R. M., Lorist, M. M., & Woldorff, M. G. (2014). Utilization of reward-prospect enhances preparatory attention and reduces stimulus conflict. Cognitive Affective & Behavioral Neuroscience, 14(2), 561-577. https://doi.org/10.3758/s13415-014-0281-z

      van Ede, F., & Nobre, A. C. (2023). Turning Attention Inside Out: How Working Memory Serves Behavior. Annu Rev Psychol, 74, 137-165. https://doi.org/10.1146/annurev-psych-021422-041757

      Vandevelde, J. R., Yang, J. W., Albrecht, S., Lam, H., Kaufmann, P., Luhmann, H. J., & Stuttgen, M. C. (2023). Layer- and cell-type-specific differences in neural activity in mouse barrel cortex during a whisker detection task. Cereb Cortex, 33(4), 1361-1382. https://doi.org/10.1093/cercor/bhac141

      Wang, L., & Krauzlis, R. J. (2018). Visual Selective Attention in Mice. Curr Biol, 28(5), 676-685 e674. https://doi.org/10.1016/j.cub.2018.01.038

      You, W. K., & Mysore, S. P. (2020). Endogenous and exogenous control of visuospatial selective attention in freely behaving mice. Nat Commun, 11(1), 1986. https://doi.org/10.1038/s41467-020-15909-2

      Zhu, P., Guan, C., Fu, Y., Shen, M., & Chen, H. (2026). Working memory encoding of attended information is adaptive to future relevance. J Exp Psychol Learn Mem Cogn. https://doi.org/10.1037/xlm0001582

    1. eLife Assessment

      This valuable study compares hippocampal-cortical functional connectivity to various other brain measures and examines their development across youth. It uses sophisticated analyses replicated in multiple datasets, but provides incomplete evidence to support the primary claim that hippocampal-cortical connectivity relates to cognitive maturation. The manuscript would benefit from a more nuanced consideration of the biological basis of some of the derived imaging measures and the limitations of the cross-sectional design. This work will be of interest to neuroimaging specialists and cognitive neuroscientists.

    2. Reviewer #1 (Public review):

      Summary:

      The authors studied the development of hippocampal connectivity gradients based on open datasets and performed correlation analyses with other MRI features as well as gene expression information from other datasets. Although the main findings are correlational and cross-sectional, the analyses are overall sophisticated and replicated in several datasets.

      Strengths:

      The hippocampus is a key region in understanding large-scale brain organization and cognition, and the authors applied advanced and suitable analytics to study its development. The paper is overall well-organized and well-written, and the findings are relevant for studying large-scale brain development.

      Weaknesses:

      While sophisticated, several of the analyses appear mainly correlational, cross-sectional, and rely on cross-dataset contextualization, which should also be stated as a limitation of the current work.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors aim to assess how the functional organisation of the hippocampus is related to the geometry and neurobiological differences of the hippocampus. In particular, the authors focus on the first three eigenvectors of hippocampal-cortical functional connectivity, based on non-linear dimensionality reduction on resting-state functional MRI data. Furthermore, the work aims to describe changes in these functional axes and their relation to other factors throughout youth and evaluate whether they are predictive of individual variations in cognition.

      Strengths:

      A major strength of this study is the attempt to replicate key findings across multiple developmental cohorts.

      Weaknesses:

      The major weaknesses of the manuscript center on gaps in technical transparency and several conceptual inaccuracies. The machine learning methodology used for cognitive prediction is scarce, leaving little means to evaluate whether the behavioral results suffer from data leakage or overfitting. The introduction sets up an oversimplified historical premise regarding the field's understanding and appreciation of hippocampal connectivity, and contains several incorrect references that throw doubt on the argumentation. Additionally, T1w/T2w signal intensity is incorrectly used as synonymous with myelin, despite gold-standard histological validation showing a non-significant correlation between T1w/T2w and myelin staining (Sandrone et al., 2013).

      Appraisal of Aims and Conclusions:

      The authors partially achieve their aims by illustrating certain age-related changes in hippocampal function; however, the correlative study design is not equipped to examine how these changes are "shaped" by geometry, myelination, or gene expression (especially the latter two). Furthermore, conclusions were often overstated based on small effect sizes.

      Context and Field Impact:

      This work adds to a growing body of literature focused on gradient-based representations of hippocampal topology. By applying these methods across a wide developmental age bracket, it provides a useful reference point for how the hippocampus and wider cortex interact during maturation. To improve utility to the neuroimaging and cognitive neuroscience communities, the nesting of subfields within the eigenvector topology should be addressed, too.

    1. eLife Assessment

      This valuable manuscript investigates how Drosophila larvae make foraging decisions in patchy environments with controlled resource density and valence; using movement tracking in bounded arenas, the authors show that larvae's patch residence time (PRT) differs depending on resource type, environmental context, and prior experience. A drift-diffusion model is used to describe patch-leaving behaviour, suggesting that an integration process may underlie stay-leave decisions during foraging. The strength of the evidence is mostly solid, but the interpretation and use of PRT needs further investigation, as PRT could be a direct effect of resource concentration on locomotion. Explicit reports of PRT statistical tests are needed for rigorous interpretation.

    2. Reviewer #1 (Public review):

      Summary:

      Mudunuri et al. investigate the foraging response of Drosophila larvae in response to patchy resources of distinct value (concentration of nutrient or valence). They show that larvae adjust their behavior according to both the quality and valence of available resources. Interestingly, previous exposure to resources of lower value increases the permanence time in resources of greater value. This suggests that larvae can value, remember and adapt their behaviour in response to previous foraging experience.

      They perform a simple integration model that recapitulates the larval behaviour.

      Strengths:

      This paper uses a very well-controlled foraging set-up where larvae are tested individually and for 3 hours, allowing for a good statistical analysis of their behaviour.

      They investigate for the first time the ability of Drosophila larvae to perceive, remember and compare the quality and valence of distinct resources. It is very exciting, as it will open up the field of foraging decision studies using the fruitfly larvae.

      Weaknesses:

      (1) Most of the analysis depends on the thresholding, but it is not clear what increasing the radius of analysis means in terms of foraging. There are two issues here:

      a) What is the behaviour of the larvae on the edges of the patch? It is obvious that the fructose or the NaCl will diffuse at the edge, so are they remaining in the proximity because they are actively feeding (exploiting) on this decaying concentration, or are they sensing the lower gradient and they are actually looking (chemosensing) for the higher concentration? The behaviour at the edge is really different (check sucrose in Wosniack et al. 2022), and there might be a way of avoiding the diffusion by actually adding a plastic ring and pouring the agar + resource in there. The effect of the ring, per se, would still have to be tested.

      b) How was the threshold selected? It is very likely that the concentration at the patch boundary will be very different for 1M and 0.1 M. Could the authors explain why they chose such a distance? What does majority of larvae mean? Is the "majority" the same for 0.1M and 1M? Is there a relationship between the threshold chosen and the diffusion of fructose and NaCl?

      (2) The word exploitation is used in the paper, but there are many instances where it is unclear whether that is the case. This should be clarified since there are no controls for exploitation.

      (3) In the experiments analysing the adaptation of foraging behaviour, it is not clear if the first and second patch means that only 2 patches were analysed per larva or the first and second in a sequence of patches visited. I think it is the second option (because of Figure S3D), but the authors should clarify this. Also, we do not know how many animals were tested. The number of data points in 4C (4G) compared to 4D (4H) seems very different.<br /> Regarding the results, which are very interesting, why aren't the larvae spending less time in the 0.1M sucrose patch after having fed on a 1M patch, while they spend more time in a 1M after a 0.1M? Could it be that the difference in residence time is correlated with their hunger rather than the comparison between conditions?

      (4) I am not an expert in this type of model, and I would appreciate it if the authors could explain how the values of the drift and leak have been fitted in Figure 5H. If possible, I would recommend adding a graph showing the parameter exploration of distinct possible combinations of values.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how Drosophila larvae make foraging decisions in patchy environments with controlled resource density and valence. Using movement tracking in bounded arenas, the authors show that larvae's patch residence time (PRT) differs depending on resource type, environmental context, and prior experience.

      The authors vary whether the environment is homogenous (all patches are equal) or heterogenous (mixed patches) and whether a higher density of the resource is appetitive (food) or aversive (salt). The most salient results are that in heterogeneous environments, larvae remain longer on higher-density patches of fructose, while they stay shorter in higher-density salt patches. The study further demonstrates that prior foraging experience influences subsequent patch residence time (PRT).

      A drift-diffusion model is used to describe patch-leaving behavior, suggesting that an integration process may underlie stay-leave decisions during foraging. Overall, the work provides a useful behavioral system for studying foraging behaviour and highlights the role of context and experience in shaping larval foraging strategies.

      Strengths:

      A major strength of the manuscript is the behavioral system. The assay is simple, well-controlled, and suitable for realistic spatial and temporal scale tracking of individual larvae. The use of non-volatile resources and embedded patches minimizes confounds from olfactory navigation and allows the authors to focus on local patch exploitation, return behavior, and experience-dependent decisions.

      The results regarding patch resident time (how long larvae stay in patches of different resource density) are convincing. In homogeneous environments, larvae spend more time on patches with a higher density of food (0.1M > 0.01M) and less time in patches with a lower density of salt (0.01M > 0.1M), indicating that their behaviour is sensitive to the valence of the resource. Further, larvae do not simply respond to current circumstances, since PRT in a given patch is sensitive to the quality of the preceding one encountered, showing some kind of memory.

      Weaknesses:

      (1) The theoretical background of the experiment, as exposed in the Introduction, is somewhat misleading. The experiment is based on patches of sufficient size for the individual larvae not to deplete them through their activity, so that the intake rate is constant while exploiting a given patch. In those circumstances, the theoretical rate-maximizing strategy would be to either reject a patch on encounter or stay in it indefinitely (until pupation). The threshold for rejection or acceptance will depend on travel time, but patch residence time would be either zero (or minimal identification time) or lifelong. In the introduction, it appears as if the system follows the classical Marginal Value Theorem assumptions as used in classical foraging theory. In that case, patch residence time is fundamentally sensitive to a decline in intake rate while in a patch. This raises questions about what factors drive patch-leaving in the present protocol. A better theoretical framework would focus on behavioural variables that can be expected to depend on the circumstances of the experiment, as discussed below.

      (2) Rather than make predictions about time in the patch, which as explained above do not reflect the present system, larval behaviour could be modelled and described as a function of observable properties such as: (a) speed of locomotion; (b) tendency to deviate from straight progress (area restricted searching); (c) probability of return after leaving a patch, possibly controlled through rea restricted searching; (d) a response to concentration gradient, since patch boundaries are probably gradual through diffusion. There is a useful literature in this regard in studies of parasitic wasps such as Venturia canescens (formerly Nemeritis canescens, see Waage 1979). Larva may respond directly to local resource concentration (see van Alphen, J. J., Bernstein, C., & Driessen, G., 2003), where higher concentration leads to increased feeding rate, reduced locomotion, and consequently results in longer time in each patch. This could still be a normative model, but based on realistic driving inputs. The dimensions of the system make it unlikely that larvae have the opportunity to adjust to travel time, or patch composition, on which classical foraging models are based. The original versions of the marginal value theorem were thought for cases where birds exploited pine cones, so that each bird had multiple encounters, and also on dung flies that mated in dung patches, which also dried out. A system with heritable optimised parameters could work for other natural systems where the parameters can be heritable, but not here.

      (3) The previous argument indicates that patch time, while it is a real quantitative consequence, is not ideal as the major dependent variable for this system. Given that the authors have the full trajectories, they could treat movement in discrete time bins and ask if the tendency to depart from linear progression (i.e. from moving straight ahead) is a function of the density of the resource. It would appear as if all the results, including return to patches (but not memory), could be explained by area-restricted searching (see Dorfman, A., Hills, T. T., & Scharf, I. (2022). A guide to area‐restricted search: a foundational foraging behaviour. Biological Reviews, 97(6), 2076-2089.). Slower movement (perhaps directly caused by eating) and more twisted progress could generate longer times in higher food densities.

      (4) The evidence for an effect of prior experience is interesting but could be strengthened. The authors state that PRT on the second patch depends on the concentration in the first patch. However, statistically significant modulation of prior experience was only found when the second food patch was richer, namely 1M fructose (Figure 4C). If the change in patch time is due to a form of learning and contrast, one might expect significantly shorter times in any second patch if the first one was richer, which is not the case. One difficulty is that the 'patchy' nature of the environment may not be evident to the larvae, because they are much smaller than the patches. From a larva's perspective, a patch is an environment, potentially suitable to remain in until pupation (which is what they ought to do in richer food patches).

      (5) The modelling section is promising but currently somewhat underdeveloped relative to the strength of the claims. The authors fit a drift-diffusion model to data and report that a drift-only model captures homogeneous environments, whereas adding a leak term improves the fit in heterogeneous environments. This provides a useful quantitative summary of behavior but the biological interpretation of the leak parameter is not clear. In addition, the valence condition was not modelled.

    4. Reviewer #3 (Public review):

      Summary:

      The work investigates how the foraging behaviour of Drosophila larvae depends on resource quality, valence, and heterogeneity in the foraging environment. A specific focus of the work was to study how foraging decisions depend on the prior experience of alternative resource patches in the same environment. Moreover, the work presents computational models (drift diffusion models) that recapitulate foraging decisions, and whose parameters appear to depend on resource quality and environment statistics, providing potential insights into the dynamics of the decision-making process.

      I am not familiar with previous literature on foraging decisions in Drosophila, but I was specifically consulted to comment on the computational modelling. Therefore, my comments will mostly focus on the modelling aspects.

      Strengths:

      In my understanding, the two strengths of the current study are that:<br /> (1) it uses non-volatile resources, providing better control of the available cues that could guide foraging decisions, and<br /> (2) it tracks foraging behaviour over an extended period of time (3h), generating a rich dataset of foraging behaviour in the same environment.

      Overall, the study appears to have been carefully conducted.

      Weaknesses:

      The computational modelling currently provides limited additional value beyond the empirical results. There are no prior hypotheses that are addressed by the computational models. Given the flexibility of DDMs, fitting foraging times is expected to be feasible. The question is whether the fits provide mechanistic insight. The main insight appears to be that describing foraging times in a homogeneous environment requires a single free parameter (drift rate), while the heterogenous environment requires a second parameter (leak). However, the effective complexity of the model is higher than the stated parameter count suggests, as each patch quality is fit with a different drift rate, which does not generalise across environments: in the heterogeneous environment, the drift rate differs substantially across fructose concentrations, whereas in the homogeneous environment, the same concentrations yield nearly identical drift rates. Counter their claims, the authors also do not systematically explore the effect of specific prior foraging experience on computational parameters, but only contrast model fits to environments with different statistics, in which prior experiences will be generally different. Overall, at the moment these modelling results have a rather descriptive character, and provide very little insight into the underlying computational principles that drive foraging decisions.

      A second weakness is that the study does not report the detailed results of the statistical tests, and it seems that the authors interpret several differences that are not marked as statistically significant in the figures. Furthermore, the model comparisons do not account for different degrees of freedom of the models, and the goodness of fit values alone are insufficient to conclude that one model is better than the other (rather than overfitting).

    1. eLife Assessment

      This useful study investigates noise-robust and energy-efficient circuit mechanisms for working memory by optimizing connectivity and reports that the resulting networks exhibit rotational dynamics and better match aspects of PFC population recording. However, the supporting evidence remains incomplete, given the restricted linear, task-specific training and analysis, and limited comparisons with other prominent models. The manuscript would be strengthened by extending the analysis to nonlinear dynamics, providing more rigorous comparisons with alternative models, and establishing a stronger link to prior theoretical and experimental work.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors address the question of working memory maintenance, starting from the experimental observation that recordings of neural activity during the delay period of working memory tasks are sometimes observed to be dynamic. They introduce a new combination of metrics (noise-robustness and energy efficiency) to quantify the performance of various network mechanisms of memory maintenance, in linear networks. They compared attractor networks, feed-forward networks, and networks trained with a loss that includes a robustness and an energy-efficiency component. They show, by plotting state-space trajectories, that networks optimized with this loss exhibit a form of rotational dynamics. They analyzed the data recorded during the delay of a working memory task in PFC, and observed state-space trajectories similar to those of the trained networks.

      The comparison with other network mechanisms is interesting in principle, but limited by the fact that only linear networks are considered. This led to counter-intuitive and misleading statements, like the fact that attractor networks are not robust to noise, or that feed-forward networks have energy consumption that is exponential in the number of neurons.

      Strengths:

      (1) The idea to use both robustness to noise and energy efficiency to assess the performance of networks on working memory tasks is interesting.

      (2) The manuscript is clearly written.

      (3) There is an interesting combination of methodologies: theory on simple models, network training, and data analysis.

      Weaknesses:

      (1) Linear networks only.

      The main feature of attractor networks is their robustness to noise, which is typically allowed by the non-linearity of neural responses. To fit their modeling framework, the authors focused only on continuous attractor neural networks (e.g., Seung 1996) and ignored point-attractor models such as the Hopfield model, which are typically used to model WM tasks, and which would presumably lead to very different results, e.g., in Figure 1D.

      The linearity assumption is also problematic for the comparison with feed-forward models. It seems that the authors obtained runaway firing rates, explaining Figure 1F middle, which are typically prevented in non-linear networks.

      The choice of parameters for the attractor network in Figure 1 is not explained. Why is t_slow = 10^4 chosen, and what does it correspond to? We expect in linear networks that activity goes back to zero or diverges as an exponential, but in principle, the time constant can be chosen to be of the same order as the time delay, with approximately linearly decreasing SNR.

      Regarding the comparison of the different mechanisms, it would have been nice to better define the notion of rotational dynamics, beyond only considering state-space analysis, which is limited to providing mechanistic interpretations.

      (2) Fixed duration of delay periods.

      I have understood that for a given network, the duration of the delay period is fixed, as opposed to a delay duration that would fluctuate from trial to trial. This would be an important assumption to relax as well, to better match common experimental paradigms, as well as to expose a fairer comparison with other network mechanisms. See Orhan and Ma (2023) for such a discussion.

      (3) Relationship with previous works

      Many other works addressed the question of dynamic firing rates during maintenance periods of WM tasks; they should be discussed and compared to the mechanism proposed here. This includes: Barak et al, Progress in Neurobio. 2013, Pereira-Obilinovic, Aljadeff, Brunel, PRX 2023, Hansel, Mato, 2013, or works pertaining to the activity-silent neural states (allowed by short-term plasticity), the framework in which the data of Panichello et al are interpreted in the original publication.

    3. Reviewer #2 (Public review):

      In this manuscript, Ritter et al. propose a model of working memory (WM) that combines feedforward and rotational dynamics. The model is discovered by optimizing a linear RNN using a loss function that encourages maximization of signal-to-noise ratio (SNR) and minimization of activation magnitude. The authors argue that the optimized model outperforms other WM models in terms of SNR and energetic efficiency, while also better replicating key features of neural responses recorded in monkey pre-frontal cortex (PFC) during a WM task. The authors also draw connections to state space models (SSM) used for other machine learning applications.

      My main issue with this manuscript is that it does not appear to convincingly demonstrate that rotational dynamics offer any advantage over purely feedforward dynamics. The authors adopt three criteria according to which they compare models:<br /> (1) SNR.<br /> (2) Energy efficiency.<br /> (3) Similarity to neural data.

      In terms of SNR, purely feedforward models seem to perform similarly to the optimized models (Figure 1). Figure 1 does seem to show that the optimized network produces responses of smaller magnitude when the number of units is large, but the authors do not explain why adding rotational dynamics would produce such a relationship. In fact, the responses that are plotted for the feedforward network in Figures 1B, 2C, and 5E look similar, if not smaller in magnitude than those of the optimized model. Lastly, while the authors claim in the body of the text that the optimized model replicates key features of monkey PFC responses better than the purely feedforward model, this is not apparent to me from the comparisons plotted in Figure 5E-J. The authors thus do not show strong evidence that the model they propose beats what they claim is an established baseline on any of the three criteria.

      Another weakness of the manuscript is that the comparison to attractor and feedforward models seems somewhat unfair. In Figure 1, the rotational model is optimized, while the parameters for the attractor and feedforward models seem to have been at least partially chosen by hand. Figure 5C again shows the three models side by side, but the fact that it compares the same network at different stages during training complicates the comparison. Instead, one should compare the rotational solution to the optimal attractor and feedforward models, respectively (obtained by constrained optimization). From looking at the flow-fields, it seems that a feedforward network with an optimized level of amplification may work just as well. On a mechanistic level, it is unclear what computational advantage rotations offer over feedforward dynamics in the WM context.

      The choice of baseline models to compare against might be questionable. The simple line attractor model by Seung et al. (1996) was initially designed to explain oculomotor integration. It is true that a line attractor has been suggested as a mechanism for working memory, e.g., in the seminal work by Machens et al (2005). However, it seems fair to say that most studies employing non-linear networks have focused on point attractors as mechanisms of working memory (e.g., Wong & Wang, 2006; Driscoll, Shenoy, Sussillo, 2024). A point attractor arguably does not suffer the SNR issues of a line attractor, because it does not lead to integration of the noise over time. However, non-trivial point attractors cannot be implemented in linear networks of the kind studied by the authors of the present study.

      The authors should expand their discussion to include other, potentially closely related work proposing rotation-like dynamics in artificial neural networks during working memory. In particular, the manuscript does not discuss Sharma, Proca, et al, ICML 2026, which describes a rotational solution to a similar WM task obtained by optimizing linear RNNs (Sharma et al., 2026, Fig. 6). Notably, Sharma et al. arrive at a similar rotational (and likely also non-normal) mechanism without using either noisy inputs or a constraint on energy efficiency. The authors of the present manuscript should discuss to what extent this finding contradicts their claim that "normative pressures on noise-robustness and energetic cost shape the complex dynamics of WM circuits." (present manuscript, Introduction). Given the obvious parallels between the two studies, a comparison between the present work and Sharma et al. (2026) would add necessary context to the Discussion.

      The authors should also clarify the significance of the "novel method for optimization of continuous-time RNNs driven by noisy inputs" (see Discussion) that the authors propose. This method is mentioned in the first line of the Discussion section but is barely discussed, let alone sufficiently explained, in the previous Sections. The only time a comparison to BPTT with a simple MSE loss is mentioned, it is stated that the two procedures produce the same results. The novel method appears to consist of a loss with two terms, the second of which is a well-known L2-penalty on unit activations (Sussillo et al., 2015). It is not clear that the method is either novel or necessary to obtain the reported results.

      Except for the fact that higher-dimensional networks also converge on rotational solutions, Figure 3 does not add much to the reader's understanding of the optimized model (except for panel F). I find the comparison to SSMs too superficial to provide real insight.

      Figure 4 claims to show that the optimized model recapitulates "a range of properties observed in prefrontal cortex and other brain areas during WM tasks" (p. 7) but does not show neural data for comparison.

    4. Reviewer #3 (Public review):

      Summary:

      The authors optimize continuous-time linear recurrent networks driven by noisy input, computing the gradient of decoding performance numerically and analytically. Optimizing for stimulus discriminability after a delay, with a penalty on firing rate, they find networks that adopt what they call high-dimensional rotational dynamics. They argue that these outperform attractor and feedforward models on noise robustness and energetic cost, and resemble state-of-the-art state-space models. They then fit a targeted dimensionality reduction model to prefrontal recordings from monkeys performing a spatial working memory task and argue that the population structure matches the rotational solution.

      Strengths:

      The evolution of the dynamics throughout learning is a nice observation, as are the analytical calculations, although I am not sure they are new since there is a fair share of work on the learning dynamics of linear networks.

      Weakness:

      I see many weaknesses. I will classify them into five groups.

      (1) Strawman comparison and no clear definition of what is rotational. The paper is centered on comparing a trained model with two models meant to represent "attractor dynamics" and non-normal dynamics. Both are picked as the weakest member of their class.

      I use quotation marks for "attractor dynamics" because I am not sure a linear system with an eigenvalue equal to zero is a representative model for the class. This is a particular linear instantiation of the line attractor from Seung 1996, but most attractor models are nonlinear and far more robust to noise, and they are robust through error correction that this linear model does not have. Even modern continuous attractors (Rivkind and Darshan) are very robust to noise through multiple mechanisms. So what the authors picked as an "attractor model" is a limited zero-eigenvalue case that, of course, will drift. "Attractor networks are highly susceptible to noise" is therefore true only of the toy they built, not of the class.

      Second, what they call a non-normal model is in fact a feedforward chain, the extreme of non-normality. There are degrees of non-normality in any matrix, and the homogeneous delay line is the corner that requires the largest firing rates. This is not representative. See Daie et al., which has a skip and recurrent structure, or Stroud, which is not a pure chain. So the feedforward chain was also picked as a strawman, chosen so that the energetic cost they then complain about is guaranteed.

      This brings me to the real problem in this section. "Rotational" is never defined. If it means complex eigenvalues, then it is a spectral property of any non-normal matrix, and "rotational versus feedforward" is not a dichotomy; it is two regions of the same continuous space of non-normal connectivity. Their own Figure 2C shows the network passing continuously through an attractor, then feedforward, then rotational during optimization. If these are points on a continuum, then "rotational dynamics is optimal" is just a statement about where the optimizer lands under this particular loss and input normalization, not the discovery of a new dynamical class. They need to define the term operationally and show the solution is qualitatively, not just quantitatively, different from non-normal feedforward. I do not think it survives that test.

      This brings me to the references.

      (2) The dynamical mechanisms of working memory have been studied for more than two decades, and I am surprised how much directly relevant work is missing. First, Druckmann and Chklovskii 2012, where a linear system produces stable encoding from oscillating modes. This is essentially their result more than a decade earlier, and it is not cited. They also miss Murray et al. on stable encoding and heterogeneous timescales in data. They oversimplify the attractor picture; for example, Pereira-Obilinovic et al. 2023 show you can have genuinely stable attractors. They do cite Daie et al., but they ignore its central claim, that non-normality is the underlying mechanism, which is more troubling than not citing it because it means they read it and did not engage. Overall, the references are idiosyncratic, missing relevant work, and not engaging the results of papers they cite.

      This brings me to the third point.

      (3) Novelty and the relationship to Stroud and Orhan. Those papers take a similar optimization approach and find that, depending on the task parameters, the optimal solution is non-normal, non-normal plus attractor, or attractor. My impression is that what this work calls rotational is just the dynamics of a strongly non-normal A, selected here by the firing-rate regularizer. They never clarify the connection with Stroud. Is the only difference the energy penalty?

      The way to settle this is quantitative, and they have the handle and do not use it: report the Henrici departure-from-normality of their optimized A and place the solution inside Stroud's regime structure.

      There is also a tension they leave implicit. In Stroud, the early loading direction is orthogonal to the late persistent readout, and that orthogonality is the source of dynamic coding. This paper's subspace alignment result (Figure 5G, H) shows exactly this early-to-late orthogonalization in both model and data, and then presents it as evidence for the rotational account and against Stroud's hybrid. You cannot reproduce a Strout's stim vs. decoder orthogonality and claim it against Strout's without doing more work.

      (4) I did not understand the SSM section, and I think it should be cut. Is this a result? Either "SSM" just means a linear dynamical system, in which case it is trivial since every linear network here, including the LMU is an SSM, or it means the network matches a fixed-connectivity model like the LMU, which it does not seem to either. So in what sense is it a result?

      (5) The data analysis is one section, and the analysis could be described as feeling somewhat like an afterthought on a very rich dataset. The coding structure they show for the rotational model also looks like the Stroud non-normal-plus-attractor model to me. They even state that the hybrid reproduces the cross-temporal subspace. What are the quantitative, cross-session metric that discriminates rotational from the non-normal-plus-attractor hybrid? Is it eyeballed trajectories?

    1. eLife Assessment

      This important study provides a detailed characterization of individual sarcomeres' contractility and of their synchrony in spontaneously beating cardiomyocytes derived from human induced pluripotent stem cells. The combination of high-resolution tracking, statistical analysis and mesoscopic modeling leads to compelling evidence that sarcomeres operate as dynamically unstable units, leading to stochastic heterogeneities in their contraction-elongation cycles depending on substrate stiffness. The work will be relevant to scientists interested in muscle biophysics, nonlinear dynamics and synchronization phenomena in biological systems.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this manuscript, the authors present comprehensive experimental observations and a theoretical framework to explain the heterogeneous behaviour of sarcomeres in cardiomyocytes. They show that a stochastic component exists in their contractile activity, which may act as a feedback mechanism regulating physiological function.

      Strengths:

      Experiments and data analysis are robust and valid. The rigorous statistical analysis and unbiased methods enable the authors to draw well-supported conclusions that go beyond the existing literature. Their outcomes inform about cellular activity at the individual level and the authors explain how the transient dynamics of single sarcomeres are governed by a force-velocity relationship and lead to the complex contractile patterns. The similarity of the results to the study cited in [24] demonstrates the validity of the in vitro setup for answering these questions and the feasibility of such in-vitro systems to extend our knowledge of out-of-equilibrium dynamics in cardiac cells.

      Very interesting the suggestion that the interplay between intrinsic fluctuations and the dynamic instability are part of a feedback mechanism for maintaining structural and functional homeostasis.

      The addition of the theoretical model and the new text of the manuscript improves the clarity of the study.

    3. Reviewer #2 (Public review):

      Summary:

      Sarcomeres, the contractile units of skeletal and cardiac muscle, contract in a concerted fashion to power myofibril and thus muscle fiber contraction.

      Muscle fiber contraction depends on the stiffness of the elastic substrate of the cell, yet it is not known how this dependence emerges from the collective dynamics of sarcomeres. Here, the authors analyze contraction time series of individual sarcomeres using live imaging of fluorescently labeled cardiomyocytes cultured on elastic substrates of different stiffness. They find that a reduced collective contractility of muscle fibers on unphysiologically stiff substrates is partially explained by a lack of synchronization in the contraction of individual sarcomeres.

      This lack of synchronization is at least partially stochastic, consistent with the notion of a tug-of-war between sarcomeres on stiff sarcomeres. A particular irregularity of sarcomere contraction cycles is 'popping', the extension of sarcomers beyond their rest length. The statistics of 'popping' suggest that this is a purely random process.

      Strengths:

      This study thus marks an important shift of perspective from whole-cell analysis towards an understanding the collective dynamics of coupled stochastic sarcomeres.

    4. Reviewer #3 (Public review):

      The manuscript of Haertter and coworkers studied the variation of the length of a single sarcomere and the response of microfibrils made by sarcomeres of cardiomyocytes on soft gel substrates of varying stiffness.

      The measurements at the level of a single sarcomere are an important new result of this manuscript. They are done by combining the labeling of the sarcomeres z line using genetic manipulation and a sophisticated tracking program using machine learning. This single sarcomere analysis shows strong heterogeneities of the sarcomeres that can show fast oscillations not synchronized with the average behavior of the cell and what the authors call popping events which are large amplitude oscillations. Another important result is the fact that cardiomyocyte contractility decreases with the substrate stiffness, although the properties of single sarcomeres do not seem to depend on substrate stiffness.

      The authors suggest that the cardiomyocyte cell behavior is dominated by sarcomere heterogeneity. They show that the heterogeneity between sarcomere is stochastic and that the contribution of static heterogeneity (such as composition differences between sarcomeres) is small.

      Strengths:

      All the results are, to my knowledge, new and original. The authors also made a theoretical model where each sarcomere is described by a Langevin equation based on a non-linear coupling between force and velocity of the sarcomeres. This model accounts well for the experimental results including the observation of what the authors call popping events.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors present comprehensive experimental observations and a theoretical framework to explain the heterogeneous behaviour of sarcomeres in cardiomyocytes. They show that a stochastic component exists in their contractile activity, which may act as a feedback mechanism regulating physiological function.

      Strengths:

      Experiments and data analysis are robust and valid. The rigorous statistical analysis and unbiased methods enable the authors to draw well-supported conclusions that go beyond the existing literature. Their outcomes inform about cellular activity at the individual level and the authors explain how the transient dynamics of single sarcomeres are governed by a force-velocity relationship and lead to the complex contractile patterns. The similarity of the results to the study cited in [24] demonstrates the validity of the in vitro setup for answering these questions and the feasibility of such in-vitro systems to extend our knowledge of out-of-equilibrium dynamics in cardiac cells.

      Very interesting the suggestion that the interplay between intrinsic fluctuations and the dynamic instability are part of a feedback mechanism for maintaining structural and functional homeostasis.

      The addition of the theoretical model and the new text of the manuscript improves the clarity of the study.

      Reviewer #2 (Public review):

      Summary:

      Sarcomeres, the contractile units of skeletal and cardiac muscle, contract in a concerted fashion to power myofibril and thus muscle fiber contraction.

      Muscle fiber contraction depends on the stiffness of the elastic substrate of the cell, yet it is not known how this dependence emerges from the collective dynamics of sarcomeres. Here, the authors analyze contraction time series of individual sarcomeres using live imaging of fluorescently labeled cardiomyocytes cultured on elastic substrates of different stiffness. They find that a reduced collective contractility of muscle fibers on unphysiologically stiff substrates is partially explained by a lack of synchronization in the contraction of individual sarcomeres.

      This lack of synchronization is at least partially stochastic, consistent with the notion of a tug-of-war between sarcomeres on stiff sarcomeres. A particular irregularity of sarcomere contraction cycles is 'popping', the extension of sarcomers beyond their rest length. The statistics of 'popping' suggest that this is a purely random process.

      Strengths:

      This study thus marks an important shift of perspective from whole-cell analysis towards an understanding the collective dynamics of coupled, stochastic sarcomeres.

      Reviewer #3 (Public review):

      The manuscript of Haertter and coworkers studied the variation of the length of a single sarcomere and the response of microfibrils made by sarcomeres of cardiomyocytes on soft gel substrates of varying stiffness.

      The measurements at the level of a single sarcomere are an important new result of this manuscript. They are done by combining the labeling of the sarcomeres z line using genetic manipulation and a sophisticated tracking program using machine learning. This single sarcomere analysis shows strong heterogeneities of the sarcomeres that can show fast oscillations not synchronized with the average behavior of the cell and what the authors call popping eveents which are large amplitude oscillations. Another important result is the fact that cardiomyocyte contractility decreases with the substrate stiffness, although the properties of single sarcomeres do not seem to depend on substrate stiffness.

      The authors suggest that the cardiomyocyte cell behavior is dominated by sarcomere heterogeneity. They show that the heterogeneity between sarcomere is stochastic and that the contribution of static heterogeneity (such as composition differences between sarcomeres) is small.

      Strengths:

      All the results are, to my knowledge, new and original. The authors also made a theoretical model where each sarcomere is described by a Langevin equation based on a non-linear coupling between force and velocity of the sarcomeres. This model accounts well for the experimental results including the observation of what the authors call popping events.

      We thank you and the reviewers for the positive evaluation of our revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Origin of the 3-Hz oscillation and required model extension. These oscillations are reproduced by our model, and their origin is already discussed in the manuscript (see lines 403–406).

      (2) Inclusion of all 5085 LOIs vs. the selected 2321. We have expanded the explanation of the LOI selection criteria in the manuscript and clarified that the main conclusions are not sensitive to this choice (lines 161-166)

      (3) Fig. 3G caption — popping rate. The caption has been updated to clarify the units and normalization. 

      (4) Fig. 4G — "Length x" vs. ΔL. Notation corrected for consistency.

      (5) Fig. 4G — gray data points. Confirmed: these represent the mean, and the caption has been updated accordingly.

      (6) Relation of k_l to the true substrate stiffness. We have added the following clarification: "The model evaluation compared the distributions of sarcomere length changes and velocities from simulations with representative experimental LOIs from substrates (5, 15, and 85 kPa, mapped to k_l = 0.5, 1.5 and 8.5 in our 1-D model; k_l is unitless, so only the ratios between values are meaningful — rescaling k_l leaves model output unchanged under correspondingly rescaled parameters) covering the full range of mechanical loads." (lines 365-369)

      (7) Could a simpler model fit the data? The cubic polynomial in Eq. (3) was deliberately chosen as a generalist ansatz rather than imposed: its coefficients were obtained by data-driven inference via Differential Evolution, and if lower-order terms within this family had sufficed, the higher-order coefficients would have been driven toward zero. The inferred nonmonotonic force–velocity relation has two extrema separated by an unstable negative-slope branch, which sets a lower bound on the polynomial order — a linear F–v is monotonic and a quadratic admits only a single extremum, so cubic is the minimum polynomial order capable of producing the observed shape. Furthermore, the qualitative phenomena we report — popping events, dynamic instability, and stochastic heterogeneity — cannot arise from any monotonic force–velocity relation, as discussed in the section on the non-monotonic instability. With 10 parameters covering complex contractile dynamics at the individual sarcomere and myofibril level across different substrate stiffnesses, the present model is parsimonious within the family of polynomial force–velocity ansätze; we have not exhaustively searched alternative non-polynomial functional families, but any such alternative would still need to reproduce the same non-monotonic shape that the data require.

      (8) Lines 497–507 in the Discussion. On reflection, we feel these lines provide useful context for the broader interpretation and would prefer to retain them.

      (9) Line 331 — motivation of Eq. (3). We have added citations to prior work motivating this form of the equation for the broader readership.

      (10) Line 427 — "scaled". Corrected.

      Reviewer #3 (Recommendations for the authors):

      We thank the reviewer for the recommendation of a theoretical appendix. The full model code, with the formulation and implementation documented in detail, is publicly available in our GitHub repository accompanying the paper, which we believe provides a complete reference for readers wishing to explore the model further. We therefore feel an additional appendix is not necessary within the scope of this revision.

    1. eLife Assessment

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells and trachea. They show that the KO mice exhibit severe hydrocephalus due to mislocated basal bodies and impaired ciliary beating. The findings are valuable with implications in the subfield of cell biology. The evidence is solid in that the methods, data and analyses largely support the claims with only a few remaining weaknesses.

    2. Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights, regarding the function of the CCP5 N-domain.

      Comments on revised version.

      The authors have appropriately revised the manuscript in response to most of my comments.

    3. Reviewer #2 (Public review):

      Summary:

      This study analyzed consequences of Agbl5 mutation on ependymal cells development and function. Authors first characterize their mutant mouse line reporting a reduced lifespan and severe hydrocephalus. Next, they report defect in ependymal cell cilia number and motility. They provide evidence for impaired basal bodies organisation, cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicate Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype are incomplete:

      Previous comment: Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting that the author checks whether the subapical network of microtubule is glutamylated or not during ependymal cells differentiation and how this network is affected in their mutants.

      Although authors now provide images of glutamylation in figure S8 their conclusion claiming that GT335 signal is increased in cilia of Agbl5M1/M1 mutant is not supported convincingly by those pictures. Quantification would be needed.

    4. Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele by extending the deletion to the N-terminus of CCP5 to investigate its function in mouse ependymal cells and trachea.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      The manuscript is well-written, and the experiments are convincing.

      Comments on revised version.

      The authors have taken all of my comments into account and have revised their manuscript to my satisfaction.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Editors for the positive assessment on our manuscript. We also thank the Reviewers for their positive remarks and constructive comments. Based on the Reviewers’ feedback, we have conducted additional experiments and provided supporting data to address Reviewers’ comments. Particularly, we provided quantitative measurement for rotational polarity of ependymal cells in Agbl5<sup>M1/M1</sup> mutants and assessed the microtubule polarization. We quantified the intensity of apical actin network in ependymal cells to strength the role of CCP5 in organizing actin network. Using scanning electron microscopy, we demonstrated the affected polarity of trachea multicilia in Agbl5<sup>M1/M1</sup>. We co-immunostained ependymal cilia with GT335 and acetylated tubulin to address the effects on their length in cilia in the mutant. We assessed the presence and length of primary cilia in ependymal cell progenitors to identify their potential contribution to the defective polarity in Agbl5<sup>M1/M1</sup> ependymal cells. We feel that these revisions have much strengthened this MS.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights.

      We thank the Reviewer’s positive comments. To address the molecular insights of the dysregulated planar cell polarity (PCP) in Agbl5<sup>M1/M1</sup> ependyma, we have conducted additional experiments to assess the microtubule polarization in ependymal cells (Figure 7O-P). We quantified the intensity of actin networks around BB patches to better understand how it is affected in the ependyma of the mutants and contributes to the dispersion of BBs (Figure 4M-N), (Please see Recommendations for the authors).

      We also assessed trachea multicilia in Agbl5<sup>M1/M1</sup> mutants using SEM and found that the polarity of trachea multicilia was affected as well (Figure S2).

      Reviewer #2 (Public review):

      Summary:

      This study analyzed the consequences of Agbl5 mutation on ependymal cell development and function. The authors first characterize their mutant mouse line reporting a reduced lifespand and severe hydrocephalus. Next, they report a defect in ependymal cell cilia number and motility. They provide evidence for impaired basal body organisation and cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicates Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype is incomplete:

      We thank the Reviewer’s constructive comments. We have performed additional quantitative analysis of the phenotypes in Agbl5<sup>M1/M1</sup> that we feel strengthen this study.

      Figure 3G - the sequence from the movie is not really informative. Providing beating frequencies as quantification of the data would be more informative.

      We have provided the beating frequency as well as the mean vector length of cilia beating directions (that reflects the coordination of cilia) in Figure 3H and 3I respectively in the revised manuscript.

      Figure 3 - the quantification of actin network would strengthen the message.

      We agree with the Reviewers. We have quantified the total intensity of actin around BBs and the actin intensity normalized to signals of the BB marker (CEP164). The data have been provided in Figure 4M and 4N respectively. The quantitative analysis showed that both the total intensity of apical actin network and the intensity of F-actin per BB are reduced in Agbl5<sup>M1/M1</sup> ependymal cells compared to that in wild-type mice, suggesting that CCP5 is involved in organizing actin network around BB. This analysis certainly improves the clarity of this message.

      Lines 219 -220 - the authors conclude «Taken together, in Agbl5M1/M1 ependymal cells, the expression of genes promoting multiciliogenesis were not impaired but certain proteins associated with differentiated ependymal cells are not properly expressed». However, they do not assess gene but protein expression (IF). In addition, their quantification shows differences in the number of FoxJ1 positive cells which indeed is an impaired expression.

      We will clarify this statement and emphasize the number of FoxJ1-positive cells.

      Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting for the authors to check whether the subapical network of microtubules is glutamylated or not during ependymal cell differentiation and how this network is affected in their mutants.

      We thank the Reviewer’s constructive comments. We conducted an immunostaining on whole-mount lateral walls of lateral ventricles for GT335 and Centrin1, the position of the latter being used to localize the subapical layer. While the GT335 signal in multicilia is increased in Agbl5<sup>M1/M1</sup> ependyma (Figure S8E), its signals underneath BBs are not much different between the mutant and wild-type (Please see Figure S8C, D, G, H).

      Showing the data mentioned in the discussion on Cep110 would be a nice addition to the paper.

      These data have been provided in Supplementary Figure S9.

      Line 354: "The latter serves as a component of tissue polarity that is required for asymmetric PCP protein localization in each cell (Boutin et al., 2014; Vladar et al., 2012)." The cited reference did not demonstrate that this microtubule network is required for asymmetric PCP localization.

      We thank the Reviewer for critical reading. The cited reference (Bountin et al., 2014) has been removed.

      Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      Weaknesses:

      The manuscript is well-written but lacks specific interpretations of the results presented. Further experiments are needed to be fully convincing.

      We thank the Reviewer’s comments. We have performed further analysis and conducted additional experiments to strengthen this study.

      (1) We have quantified the intensity of actin staining around BB patches and its intensity relative to the number of BBs to assess to which extent the actin networks in Agbl5<sup>M1/M1</sup> ependymal cells are affected (please refer to the above response to the comments of Reviewer 2#). The results were shown in Figure 4M-N.

      (2) We Co-stained tdTomato with an ependymal cell-specific markers to strengthen the expression of Agbl5 in ependymal cells (please see Figure 6C-E).

      (3) We have conducted co-immunostaining of GT335 and Ac-Tub and compared the length of their signals in ependymal multicilia between WT and Agbl5<sup>M1/M1</sup> mice (please see Figure 6O, P, R, S).

      (4) We quantified the area of ependymal cells in the wild-type and Agbl5<sup>M1/M1</sup> mice. Indeed, the area of ependymal cells is increased in the mutants. However, the primary cilia are present in the ependymal cell progenitors of Agbl5<sup>M1/M1</sup> mice and have similar length with that in the wild-type (Please see Figure 7M, N and our response to this point below).

      (5) We performed additional analysis to address the affected rotational polarity in the Agbl5<sup>M1/M1</sup> mutant mice (please see Figure 3I, Figure 7E).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors showed that the actin networks were severely affected, leading to impaired stability of basal bodies and that the intensity and length of acetylated tubulin signal in the multicilia were dramatically reduced in AGBL5M1/M1mutant mice (Figures 3 and 5). Data also suggested the dysregulation of planar cell polarity. Are expression and localization of other planar cell polarity proteins such as tyrosinated tubulin and Fzd6 affected in mutant mice?

      We thank the Reviewer’s recommendations. We have assessed the expression of tyrosinated tubulins and found they are similarly polarized in ependymal cells from wild-type and Agbl5<sup>M1/M1</sup> mice. The results are presented in Figure 7O, P in the revised MS. We also tried to assess the expression of Fzd6. However, with the antibody we tested, Fzd6 signals were not convincing. Therefore, we prefer to not showing the results and drawing a conclusion on it.

      (2) The phenotype of multiciliated cells in tracheas should also be examined in mutant mice. It is important to elucidate whether AGBL5 commonly functions in multiciliated cells of other organs.

      We thank the Reviewer’s suggestion. We have assessed the multicilia in the tracheas of P30 mice using scanning electron microscopy. Indeed, unlike the multicilia in wild-type mice that orientate to the same direction, those in the tracheas of Agbl5<sup>M1/M1</sup> mice often radiate to different directions in individual cells (Figure S2). Therefore, Agbl5 appears commonly involved in the alignment of multicilia.

      (3) According to Figure 1B, AGBL5 is highly expressed in the brain. Which cells in the brain express it besides ependymal cells?

      Based on the localization of tdTomato tracer engineered in Agbl5 mutant alleles (Figure 5B), Agbl5 is broadly expressed in the brain, including most if not all neurons, but its expression is much weaker in the subventricular zone (Please see Figure 5B). We clarified this in the revised MS.

      (4) From a mechanistic point of view, it is necessary to identify binding proteins with the N-domain of AGBL5 and perform functional analyses.

      We agree with the Reviewer. We feel that identification of the binding partners of CCP5 N-domain and functional analysis may be more suitable to go along with other mechanistic analysis on the function of CCP5 in ependymal cell polarities in our future study.

      Reviewer #2 (Recommendations for the authors):

      (1) Movie 3: The authors could comment on beating direction that seems impaired at the cell scale here, analysis of rotational polarity would be a plus.

      We thank the reviewer’s recommendation. We have analyzed the beating directions of cilia in individual cells and presented their consistency in each cell using mean vector length. These results indeed demonstrated defective rotational polarity in the cell level in Agbl5<sup>M1/M1</sup> mice (please refer to Figure 3I). We also analyzed the beating directions of ependymal multicilia in earlier stage in tissue level (Figure 7E). The mean vector length of cilia beating direction in Agbl5<sup>M1/M1</sup> mice is significantly reduced compared to that in wild-type, suggesting an aberrant rotational polarity in the tissue level in the mutant (Figure 7E).

      (2) Line 166 : ref to Werner et al., 2011 is not correct (no ependymal cells in that paper).

      We thank the reviewer’s critical reading. This reference has been removed.

      (3) Figure S4: B and D look similar picture to me same for C and F.

      We apologize for using the wrong images in this Figure. It has been corrected (Revised Figure S5).

      (4) Line 328: "Therefore, CCP5 apparently contributes to the establishment of both translational and tissue polarities in ependymal cells." Should be rephrased since translational polarity is also a tissue-level parameter which is the coordinated positioning of the ciliary patch. Cf Mirzadeh et al., 2010; Boutin et al., 2014.

      We thank the Reviewer’s comments. The sentence has been rephrased. This concept has been clarified where else needed in the revised manuscript. 

      (5) Line 348: "Planar cell polarity (PCP) pathway is essential for the establishment of rotational and tissue polarities in ependymal cells" Rotational polarity also has a tissular component (ie coordination of beating direction across tissue which is reflected by coordination of basal body polarities across tissue).

      We thank the Reviewer’s comments. We have clarified this point in the revised MS.

      (6) Incomplete bibliography citation (ie Walentek et al. without date).

      We thank the Reviewer’s critical reading. This bibliography citation has been fixed.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 3: The authors assert that the mutant's apical actin networks are significantly disrupted. However, the cell shown in Figure 3Q-R exhibits less compact centrioles than the controls, which could account for the reduction in phalloidin staining. Because centriole dispersion is variable in the mutant, quantifying actin staining in representative cells would be necessary to support such a statement.

      We thank the Reviewer’s comments. To address this concern, we have quantified the total intensity of actin network around BBs as well as the intensity of F-actin signals normalized to the level of immunosignals of BBs ((revised Figure 4M, N) please also refer to our response to Reviewer 1#). The results indicated the intensity of actin signal per BB is reduced in the mutant compared to that of wild-type mice. We feel that this analysis strengthened our statement.

      (2) Figures S3 and 4A-B show that the authors examine tdT expression to show that Agbl5 is expressed in ependymal cells but not in the SVZ. However, the tdT signal intensity is very low, and cells are very dense in this brain region. Double staining with specific markers of ependymal and/or SVZ cells would help convince readers that tdT is not expressed in SVZ cells.

      We agree with the Reviewer that the intensity of tdT signal is low, but broadly detectable in brain. Compared with its expression in ependymal cells, that in SVZ is much lower if any (Figure 4B’). To further confirm the identity of tdT-positive cells along the surface of ventricles, we have co-stained the brain sections of Agbl5<sup>WT/M1</sup> mice for tdT and S100b, a marker of mature ependymal cells (Figure 5C-E). The signal of tdt is colocalized with that of S100b and is much lower in cell layers next to S100b-positive cells.

      (3) Figure 4C-D and S4: The authors demonstrate that the number of FoxJ1+ cells per section increases at P7 (4C-E), while the number of S100β+ cells per mm decreases. Quantifications should be carried out in a similar manner to ensure comparability (number of positive cells per mm). Additionally, it remains unclear how to interpret these results, as S100β and FoxJ1 are two markers of differentiated cells, yet they exhibit opposite trends compared to controls. Is this a direct or indirect effect of Agbl5 mutation? The increase in the number of FoxJ1+ cells is particularly surprising given that the number of GT335 multicilia per mm remains unchanged (Figure 5).

      We agree with the Reviewer that quantifications should be carried out in a similar manner. In the revised MS, the quantification of Foxj1-positive cells is presented in number per mm (Figure 5I). To be noted, the expression of Foxj1 was assessed at P7 when ependymal cells are differentiating. while the expression of S100β was assessed at P17 when ependymal cells are supposed to be fully mature. Although S100b is used as a marker of mature ependymal cells, given its unclear function, we removed the results of S100b-positiving cell counting to avoid confusion in the revised manuscript.

      (4) Figure 5: In this figure, the authors analyze the labeling obtained with GT335, Acetylated Tubulin, and Arl13b antibodies. They show that the area of the cilium labeled by GT335 has increased, while the area labeled by the Acetylated Tubulin antibody has decreased in the knockout (KO) compared to the control. However, the length of the cilia observed through labeling with the Arl13b antibody remains unchanged. These observations are intriguing, but the low-magnification images in Figure 4 do not allow for the differences in ciliary axoneme labeling to be seen. Double GT335/AcTub labeling and higher magnifications are necessary for improved visualization of the differences in labeling along the axonemes.

      We thank the Reviewer comments. We have co-stained the cilia with GT335 and Ac-Tub antibodies, re-quantified cilia length labeled with respective antibodies and provided high magnification images. Please see the revised Figure 6O,P,R,S.

      (5) Figure 6: An analysis of ciliary beats using a high-speed camera shows no difference in ciliary beat frequency between the control and KO groups. At least, 3 animals should be analyzed. According to Figure 5, these findings indicate that the decrease in ciliary acetylation and the increase in ciliary glutamylation do not affect the beat frequency; instead, they disrupt the orientation of the beats. While these results are intriguing, they require further confirmation. Analyzing ciliary beats with a high-speed camera is informative, but at least three animals per genotype should be examined to ensure rigor. Furthermore, if the coordination of ciliary beats is impaired within the cells, this should be validated by double-labeling centrioles and basal feet to demonstrate that the orientation of cilia within the cells is abnormal.

      We thank the Reviewer’s comments. Sections shown in Figure 5 (currently Figure 6) are from P7 mice, while the ciliary beating analysis shown in Figure 6 (currently Figure 7) is from P15 mice. As the PTM changes in cilia were also observed in Agbl5<sup>M2/M2</sup>, we don’t think this is the cause that disrupts the orientation of the beats. The rotational polarity of Agbl5<sup>M1/M1</sup> ependymal cells is affected. Please refer to the analysis in Figure 3I and Figure 7E in the revised manuscript.

      (6) Figure 6F-G: β-Catenin labeling reveals cells of varying sizes in the KO. This phenotype is typical of ciliary mutants that lack primary cilia (Mirzadeh et al., 2010). Hence, it is essential to examine the mutation's impact on the presence, length, and positioning of the primary cilium in ependymal cell progenitors.

      We thank the Reviewer’s constructive comments. We assessed the area of ependymal cells labeled with β-Catenin. Indeed, the ependymal cells in the mutant showed larger area than that of wild-type. The ratio of the area of BB patch over that of cell surface is reduced (please see Figure 7O, P in the revised manuscript). However, primary cilia are present in ependymal cell progenitors in the mutant and exhibit comparable length with those in the wild-type (Figure S8). Due to some technique problems, we were unable to get convincing results from whole-mount ventricle walls for the primary cilium positioning at this time. We speculate that the localization of certain sensory proteins in primary cilia or the positioning of primary cilia might be affected in Agbl5<sup>M1/M1</sup> mice. We discussed this possibility and will certainly systemically assess this intriguing aspect in our future investigation.

      (7) Given the regular beating frequency in the KO at P15, how do the authors explain the complete absence of ciliary beating in the adult? How many animals were analyzed? One would expect ciliary beating to remain unaffected as it was at P15 unless the cilia structure was specifically altered at the adult stage. Is that the case?

      We thank the Reviewer’s critical questions. We do think that the ciliary structure of Agbl5<sup>M1/M1</sup> ependymal cells is likely altered during aging. Given that only Agbl5<sup>M1/M1</sup> but not Agbl5<sup>M2/M2</sup> mice develop hydrocephalus, we speculate the N-domain of CCP5 may contribute to the integrity of ependymal multicilia. We have added this in the Discussion section. For each genotype, 2 mice were analyzed.

      (8) Line 264 of the manuscript: replace intercellular with intracellular.

      It has been revised.

      (9) Indicate the number of animals analyzed in each experiment

      It has been included in figure legends.

    1. eLife Assessment

      This paper addresses a valuable research question on the modest heritability of the brain's response to movie watching, and how heritability varies under different parameters such as regional spatial hyperalignment and BOLD frequency bands. The topic of this paper is of interest to fMRI methodological experts, and potentially to a broader cognitive neuroscience audience, and those with an interest in understanding the heritable sources of individual differences in brain function. Although some of the conclusions could be strengthened by future cross validation studies in independent and larger family-based samples, and through complementary twin/family and SNP-based models, taken altogether, the analyses and results provide convincing evidence for the overall conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales, and more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor our topographic differences and found the relationship between heritability and neural time scales very interesting. The writing is clear and the results are compelling. In general, I don't have many complaints after a couple reads through the manuscript; most of my comments below are relatively minor suggestions and points of clarification.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified:

      On page 16, you compare heritability in functional connectivity (FC) and response time series and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and time-series differences cannot, by definition). This makes me wonder how this connectivity result would change if you used intersubject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript-but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. You could even color the data points according to networks like in Figure 3C. (You also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      On page 9, if I understand correctly, you regress the vector of ISC values across parcels out of the vector of heritability values across parcels and then plot the residual heritability values. Do you center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can you explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could you go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?)

      On page 4 (line 155), you say "we shuffled dyad labels"-is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure your approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned at line 189?

      I found panel A in Figure 4 to be a little bit misleading because your parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels you're performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248-259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Comments on revised version.

      The authors have adequately addressed my previous comments. This is a strong contribution: the methods are sophisticated, the statistical treatment is rigorous, and the results are quite interesting/compelling. I'm happy to endorse the revised manuscript as a finalized version.

      Just to confirm: The subjects watched all different movies across the two days, right? For a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

    3. Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      Comments on revised version.

      The whole manuscript has been improved a lot, and the concerns have been clarified.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained by local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales and is more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor out topographic differences, and I found the relationship between heritability and neural time scales very interesting. The writing is clear, and the results are compelling.

      We thank Reviewer 1 for their kind words and enthusiastic support of our manuscript.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified.

      (1) On page 16, the authors compare heritability in functional connectivity (FC) and response time series, and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and timeseries differences cannot, by definition). This makes me wonder how this connectivity result would change if the authors used inter-subject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript, but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      We agree that investigating the heritability of ISFC (or stimulus-driven functional connectivity) would make for a very interesting future direction. Ultimately, we chose to analyze FC (vs. ISFC) profiles to allow for direct comparison with the sizable existing literature on the heritability of FC (such as in our Movie vs. Rest FC analysis) and decided to refrain from analyzing ISFC data in order to keep the present manuscript focused. ISFC analysis of this dataset will be a focus of future work.

      (2) The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. The authors could even color the data points according to networks, like in Figure 3C. (They also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      We thank R1 for this helpful suggestion- we originally set the y-axis limits to r = 1 in order to facilitate comparison between ISC (Fig. 1B) and FC profile (Fig. 6B) similarity, but we agree that this renders the group differences harder to discern and have updated the plot accordingly (along with thicker lines to enhance readability). We prefer to keep the line plots in the main body as they allow for direct comparison of all three groups on the same plot, but we have included the scatter plot version in Fig. S2 for those who are interested.

      (3) On page 9, if I understand correctly, the authors regress the vector of ISC values across parcels out of the vector of heritability values across parcels, and then plot the residual heritability values. Do they center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can the authors explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could they go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?).

      We indeed included an intercept in this model using MATLAB’s fitlm function. This means that the model estimates the best-fitting line of the following form: heritability<sub>i</sub>=β0+β1ISC<sub>i</sub> +ε<sub>i</sub>. We agree that the interpretation of these ε<sub>i</sub> values and alternative approaches to controlling for ISC should be clarified. As such, we have added the following passages to the text:

      Methods: “Because the heritability of ISC is constrained by the degree of synchronization in a given area, we also sought to identify areas in which BOLD time courses were more/less heritable than would be expected based on ISC alone by fitting a linear model of the form heritability<sub>i</sub>=β0+β1ISC<sub>i</sub>+ε<sub>i</sub> and plotting the residuals. Regarding alternative approaches to controlling for ISC, although the heritability model introduced by Ge et al. allows for the inclusion of covariates defined at the subject level (e.g., age), it does not allow for covariates that are defined at the dyad level (e.g., pairwise ISC).”

      Results: “Here, negative values in the residual map indicate parcels where heritability is lower than expected based on ISC, while positive values indicate higher-than expected heritability.”

      (4) On page 4 (line 155), the authors say "we shuffled dyad labels"- is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure their approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned in line 189?

      Briefly, shuffling the kinship matrix involved permuting the rows and columns of the matrix in the same manner (also known as the quadratic assignment procedure), whereas shuffling the dyad labels involved random permutations of the three group labels (MZ, DZ, unrelated), which could not be done through matrix operations as the age- and gender matching precluded the use of a complete similarity matrix. However, given concerns raised by Reviewer 2, we have removed our significance claims from this (and similar) sections, which we discuss in more detail in response to Reviewer 2’s weakness A.

      (5) I found panel A in Figure 4 to be a little bit misleading because their parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels they are performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      We agree that our efforts to present a simplified depiction of hyperalignment may mislead less familiar readers and have amended Fig. 4A according to this suggestion. We have also added text to the methods section (below) to clarify that the outputs of hyperalignment are time series that reflect linear combinations of other voxels’ time series from that parcel.

      “This approach independently transforms each subject's data within discrete anatomical parcels into the common space, yielding functionally aligned vertex time series that are calculated as weighted linear combinations of the original time series from all other vertices within that same parcel for that subject.”

      (6) I believe the subjects watched all different movies across the two days, however, for a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

      We agree that this would be helpful and have added the following text to the relevant sections:

      “All clips were only viewed once by each subject, with the exception of the brief montage which was included at the end of each of the four runs for test-retest purposes.”

      “To characterize the heritability of brain responses to complex stimuli, we used 7T fMRI data from 178 HCP Young Adult subjects acquired across two days (using two largely non-overlapping sets of movie stimuli, see Methods)…”

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to estimate the heritability of brain activity evoked from a naturalistic fMRI paradigm. No new data were collected; the authors analyzed the publicly available and well-known data from the Human Connectome Project. The paper has 3 main pieces, as described in the Abstract:

      (1) Heritability of movie-evoked brain activity and connectivity patterns across the cortex.

      (2) Decomposition of this heritability into genetic similarity in "where" vs. "how" sensory information is processed.

      (3) Heritability of brain activity patterns, as partially explained by the heritability of neural timescales.

      Strengths:

      The authors investigate a very relevant topic that concerns how heritable patterns of brain activity among individuals subjected to the same kind of naturalistic stimulation are. Notably, the authors complement their analysis of movie-watching data with resting-state data.

      Weaknesses:

      The paper has numerous problems, most of which stem from the statistical analyses. I also note the lack of mapping between the subsections within the Methods section and the subsections within the Results section. We can only assess results after understanding and confirming the methods are valid; here, however, Methods and Results, as written, are not aligned, so we can't always be sure which results are coming from which analysis.

      (A) Intersubject correlation (ISC) (section that starts from line 143): "We used nonparametric permutation testing to quantify average differences in ISC for each parcel in the Schaefer 400 atlas for each day of data collection across three groups: MZ dyads, DZ dyads, and unrelated (UR) dyads, where all UR dyads were matched for gender and age in years." ... "some participants contributed to ISC values for multiple dyads (thus violating independence assumptions)"

      This is an indirect attempt to demonstrate heritability. And it's also incorrect since, as the authors themselves point out, some subjects contribute to more than one dyad.

      Permutation tests don't quantify "average differences", they provide a measure of evidence about whether differences observed are sufficient to reject a hypothesis of no difference.

      Matching subjects is also incorrect as it artificially alters the sample; covarying for age and sex, as done in standard analyses of heritability, would have been appropriate.

      It isn't clear why the authors went through the trouble of implementing their own nonparametric test if HCP recommends using PALM, which already contains the validated and documented methods for permutation tests developed precisely for HCP data.

      The results from this analysis, in their current form, are likely incorrect.

      We appreciate that permutation tests do not quantify average differences and intended to write “We used non-parametric permutation testing to quantify [the significance of] average differences…”. Our intention with this analysis was not to demonstrate heritability, but rather to quantify group differences in ISC in a manner that is interpretable for readers who are unfamiliar with h<sup>2</sup> (e.g., “identical twins’ BOLD time courses were 59% more similar than those from pairs of unrelated individuals”) and motivate the formal heritability analysis used later in the paper. Indeed, all of the heritability analyses in this paper leveraged a validated multidimensional heritability method first introduced by Ge et al. (2016) and used by many other investigators since then. Furthermore, we covaried for age and sex at the subject level in all our heritability analyses, and always tested the significance of these heritability values using a validated permutation procedure (the quadratic assignment procedure; Hubert & Schultz, 1976) that respects the non-independence of dyadic data.

      Regarding the shuffling procedure used for Figure 1, while PALM is the standard for univariate, subject-level GLMs in the HCP pipeline and can accommodate nested designs (i.e., subjects within families), it is not designed to handle the unique relational dependencies of dyadic ISC analysis (i.e., the same subject contributing to multiple dyads). Although the element-wise resampling approach was the most appropriate approach available, it is known to inflate the false positive rate (Chen et al., 2016; doi:10.1016/j.neuroimage.2016.05.023); given that this analysis was simply meant to motivate our later hypothesis testing heritability analyses, we have removed significance claims from this section of the manuscript. Still, we emphasize that this has no bearing on the validity of our conclusions which were supported by our formal heritability analyses; throughout our paper we have correctly used the appropriate methods to back the stated claims.

      (B) Functional connectivity (FC) (section that starts from line 159): Here the authors compute two 400x400 FC matrix for each subject, one for rest, one for movie-watching, then correlate the correlations within each dyad, then compared the average correlation of correlations for MZ, DZ, and UR. In addition to the same problems as the previous analysis, here it is not clear what is meant by "averaging correlations [...] within a network combination". What is a "network combination"? Further, to average correlations, they need to be r-to-z transformed first. As with the above, the results from this analysis in its current form are likely incorrect.

      We regret that R2 had difficulty understanding our analysis and have added the following text to the relevant Methods section to clarify our approach:

      “For example, there are 16 parcels in the Kong et al. Auditory network and 17 parcels in the Language network, so the FC profile for a given subject’s Auditory-Language network combination consists of the (16 * 17 =) 272 correlation coefficients between all unique pairs of one parcel from each network.”

      As we stated in the previous Methods paragraph, “All Pearson r values in this and all other analyses were Fisher z-transformed before averaging (and converted back to Pearson r for visualization)”. Thus, contrary to the reviewer’s assertion, these analyses were performed correctly. Once again, we emphasize that this analysis was not intended to demonstrate heritability, but rather to describe group differences in FC in familiar units.

      (C) ISC and FC profile heritability analyses (section that starts from line 175): Here, the authors use first a valid method remarkably similar to the old Haseman-Elston approach to compute heritability, complemented by a permutation test. That is fine. But then they proceed with two novel, ill-described, and likely invalid methods to (1) "compare the heritability of movie and rest FC profiles" and (2) to "determine the sample size necessary for stable multidimensional heritability results". For (1), they permute, seemingly under the alternative, rest and movie-watching timeseries, and (2), by dropping subjects and estimating changes in the distribution.

      The (1) might be correct, but there are items that are not clearly described, so the reader cannot be sure of what was done. What are the "153 unique network combinations"? Why do the authors separate by day here, whereas the previous analyses concatenated both days? Were the correlations r-to-z transformed before averaging?

      The (2) is also not well described, and in any case, power can be computed analytically; it isn't clear why the authors needed to resort to this ad hoc approach, the validity of which is unknown. If the issue is the possibility that the multidimensional phenotypic correlation matrix is rank-deficient, it suffices that there are more independent measurements per subject than the number of subjects.

      Regarding (1), we have clarified in section 2.6 that the 153 unique network combinations reflect each unique pair of 17 Kong networks. All of our analyses, including this one, were performed separately for each day of data collection, as we state throughout the paper and visualize in our figures (although we acknowledge that, on some occasions, we [conservatively] performed FDR-correction on a combined set of p-values, as discussed in our response to K). Given that the null hypothesis for this analysis is that rest FC and movie FC are equally heritable, we are not sure why permuting rest and movie FC matrices would be invalid. All Pearson r values were z-transformed before averaging, as we stated in our paper.

      Regarding (2), we included this analysis in response to editorial concerns that our heritability analyses were not sufficiently powered, and we chose this approach because it serves as a simple way to demonstrate the stability of our results at various sample sizes whose validity is self-evident. Furthermore, this sort of subsampling approach has been used many times before in our field (e.g., Marek et al., 2022) and others (e.g., Manyara et al., 2024) to demonstrate the sample-size dependence and stability of statistical effects. We have added text explaining this to the relevant Methods section (2.6).

      (D) Frequency-dependent ISC heritability analysis (from line 216): Here, the authors decompose the timeseries into frequency bands, then repeat earlier analyses, thus bringing here the same earlier problems and questions of non-exchangability in the permutations given the dyads pattern, r-z transforms, and sex/age covariates.

      We did not use dyadic permutation testing for any of the frequency-dependent ISC analyses; rather, we used the jackknife SEMs to compare heritability across frequency bands and have added an explicit description of this to section 2.7. We have addressed the r-z transform and covariate concerns in previous comments.

      (E) FC strength heritability analysis (from line 236): Here, the authors use the univariate FC to compute heritability using valid and well-established methods as implemented in SOLAR. There is no "linkage" being done here (thus, the statement in line 238 is incorrect in this application. SOLAR already produces SEs, so it's unclear why the authors went out of their way to obtain jackknife estimates. If the issue is non-normality, I note that the assumption of normality is present already at the stage in which parameters themselves are estimated, not just the standard errors; for non-normal data, a rank-based inversenormal transformation could have been used. Moreover, typically, r-to-z transformed values tend to be fairly normally distributed. So, while the heritabilities might be correct, the standard errors may not be (the authors don't demonstrate that their jackknife SE estimator is valid). The comparison of h2 between dyads raises the same questions about permutations, age/sex covariates, and r-z transforms as above.

      We used jackknife SEs for these analyses to maintain consistency with the multidimensional heritability package used here, which only outputs jackknife SEs. We note that this jackknife approach (and the corresponding multidimensional heritability analysis) was detailed in prior work (Anderson et al., 2021), and that the leave-one-family-out jackknife has a long history of being used to estimate SEs in heritability studies, especially when working with smaller samples (Knapp et al., 1989). We are also not sure what “the comparison of h2 between dyads” means- heritability cannot be compared “between” dyads; rather, it is defined across dyads.

      (F) Hyperalignment (from line 245): It isn't clear at this point in the manuscript in what way hyperalignment would help to decompose heritability in "where vs. how" (from the Abstract). That information and references are only described much later, from around line 459. The description itself provides no references, and one cannot even try to reproduce what is described here in the Methods section. Regardless, it isn't entirely clear why this analysis was done: by matching functional areas, all heritabilities are going to be reduced because there will be less variance between subjects. Perhaps studying the parameters that drive the alignment (akin to what is done in tensor-based and deformation-based morphometry) could have been more informative. Plus, the alignment process itself may introduce errors, which could also reduce heritability. This could be an alternative explanation for the reduced heritability after hyperalignment and should be discussed. An investigation of hyperaligment parameters, their heritability, and their co-heritability with the BOLD-phenotypes can inform on this.

      To help set up our hyperalignment analyses, we have added text to the introduction explaining how hyperalignment would help to decompose heritability. The description in the Methods section included a reference to Bazeille et al., 2021, in which the hyperalignment method used here is discussed in detail. Still, we have added citations to additional papers (also cited in the Bazeille et al. paper, and elsewhere in our paper) in case that might be helpful. We note that it is not the case that all heritabilities were reduced by hyperalignment- as can be seen in Figs. 4D, 8A, and S15, hyperalignment did increase heritability in some voxels and network combinations. This would be expected under the alternative (albeit unlikely) hypothesis that functional topographies are not heritable, such that topographic variation between related individuals would obscure similarities in their (heritable) topography-independent brain responses. Recognizing that this alternative is unlikely, we believe the main novelty of this analysis comes from the magnitude of the hyperalignment effect (up to 40% of brain-wide heritability) and its spatial pattern (e.g., larger heritability decreases in visual vs. auditory cortex, the opposite of our NT result).

      We agree that we would see lower post-hyperalignment heritability if the alignment process itself introduced errors/noise, but this would be deeply surprising as hyperalignment increases ISC by design (and errors/noise could only decrease ISC). To demonstrate this, we have added Figure S7 which shows that (as expected) ISC across all voxels and subject pairs increases after hyperalignment (and that this increase is larger when hyperalignment is performed in larger parcels). Given that hyperalignment increased ISC, and that it is blind to twin status, we are unsure how it could have introduced errors that would have confounded this result.

      (G) Relationships between parcel area and heritability (from line 270): As under F), how much the results are distorted likely depends on the accuracy of the alignment, and the error variance (vs heritable variance) introduced by this.

      We agree that alignment accuracy could potentially impact parcel-level differences in how much heritability changes following hyperalignment, and we included the frequency dependent h<sup>2</sup><sub>residuals</sub> (controlling for differences in ISC) in Fig. 3 for this reason, as more accurate hyperalignment should result in greater increases in ISC, raising the heritability ceiling. We note that we observe similar relationships between parcel rank and frequency dependent changes in these residualized maps, suggesting that our parcel-level differences are not simply the result of better alignment in more sensory parcels.

      (H) Neural timescale analyses (from line 280): Here, a valid phenotype (NT) is assessed with statistical methods with the same limitations as those previously (exchangability of dyads, age/sex covariates, and r-z transforms). NT values are combined across space and used as covariates in "some multivariate analyses". As a reader, I really wanted to see the results related to NT, something as simple as its heritability, but these aren't clearly shown, only differences between types of dyads.

      We have addressed the exchangeability, covariates, and r-z transform comments above (in A). As we explained for our FC strength analyses, we are underpowered to evaluate the heritability of unidimensional traits (like the heritability of NT magnitude), and the heritability of a closely-related measure (BOLD turnover magnitude) has already been established in a larger sample of HCP subjects (https://doi.org/10.1152/jn.00402.2022). Still, we agree that more results related to the heritability of NTs would be of interest to our readers. As such, we have added an analysis in section 3.4 quantifying the heritability of multivariate NT topographies and used SOLAR to quantify the heritability of NT magnitudes, with the disclaimer that this and similar analyses are underpowered (hence the large difference in day 1 and day 2 heritability effect sizes). We also removed significance claims for the dyadic NT similarity analysis.

      (I) Significance testing for autocorrelated brain maps and FC matrices (from line 310): Here, the authors suddenly bring up something entirely different: reliability of heritability maps, and then never return to the topic of reliability again. As a reader, I find this confusing. In any case, analyses with BrainSMASH with well-behaved, normally distributed data are ok. Whether their data is well behaved or whether they ensured that the data would be well behaved so that BrainSMASH is valid is not described. As to why Spearman correlations are needed here, Mantel tests, or whether the 1000 "surrogate" maps are valid realizations of the data under the null, remains undemonstrated.

      We brought up reliability in this section because we show the reliability of our results across the two days of data collection several times in the paper. R2 is correct to point out that BrainSMASH was validated using normally distributed brain maps, and although some of our brain maps contain normally distributed values, others are right skewed (due largely to the fact that many voxels/parcels exhibit low ISC while visual/auditory areas have very high ISC). In preparing our original manuscript, we visualized BrainSMASH’s variogram outputs for one of the most skewed inputs (vertex-wise BOLD time course heritability) and found that the autocorrelation structures of the empirical and null maps were well-matched. We did not include this in the original manuscript as it is not commonplace in the field to report the variograms, see Author response image 1. Furthermore, our use of Spearman (vs. Pearson) correlations renders these distributional differences less relevant, as the Spearman correlation transforms all inputs to a uniform distribution. To empirically check that these distributional differences do not bias our results, we retested the significance of all brain map associations using the spin test (10.1016/j.neuroimage.2018.05.070), an alternative method that does not assume normally distributed inputs, and obtained identical p-values for all analyses (P<.001 in all cases).

      Author response image 1.

      (J) Global signal was removed, and the authors do not acknowledge that this could be a limitation in their analyses, nor offer a side analysis in which the global signal is preserved.

      Although we agree that GSR is a contentious preprocessing step for certain analyses, it has explicitly been shown to increase ISC signal-to-noise without compromising FC fingerprints (Graff et al., 10.1016/j.dcn.2022.101087), and it is uncommon to perform ISC analyses with and without GSR. Still, we have added additional text to our Methods section explaining our rationale for using GSR and that this could affect our results. We also re-ran our main analysis (BOLD time course heritability) with and without GSR and found that GSR had little impact on our results; we have included this in our manuscript as Fig. S4.

      Specifically, we see that GSR resulted in a slight increase in heritability (average Day 1 h<sup>2</sup> with/without GSR = .064/.060; Day 2: .068/.061) and almost no effect on the spatial pattern of our results (With GSR/without GSR Spearman ρ = .99, P<sub>brainSMASH</sub> < .001 on both Day 1 and Day 2).

      (K) FDR is used to control the error rate, but in many cases, as it's applied to multiple sets of p-values, the amount of false discoveries is only controlled across all tests, but not within each set. The number of errors within any set remains unknown.

      We agree that the FDR usage in our original manuscript was inconsistent, in that for two analyses we FDR-corrected p-values from the two days of data collection together (instead of correcting p-values from each day separately and reporting voxels/parcels/etc. that were significant at q<.05 on both days, as in the rest of our analyses). We note that both approaches are more conservative than reporting significant results at q<.05 separately; regardless, to maintain consistency we have updated all analyses such that FDR correction is always performed separately for each day of data collection.

      (L) Generally, when studying the heritability of a trait, the trait must be defined first. Here, multiple traits are investigated, but are never rigorously defined. Worse, the trait being analyzed changes at every turn.

      Here, we analyze the heritability of movie-evoked BOLD time courses (Figures 1-5) as well as FC profiles (Figures 6-8). We defined FC profiles in our Introduction as an individual’s pattern of pairwise FC strengths (and further detailed how we quantified FC profiles in the relevant Methods section), and believe that “BOLD time course” is a well understood phrase in the field and does not need to be further defined. We also used hyperalignment to decompose the heritability of these traits into topography-dependent and independent portions, and (new to this version) also explicitly quantify the heritability of neural timescales, which we defined as the AUC of the ACF until the first negative ACF value in both the relevant Results and Methods sections.

      To make this clearer, we have modified the last paragraph of our Introduction to begin with:

      In the present work, we address these questions by analyzing 7T fMRI recordings of a twin sample acquired by the Human Connectome Project (Van Essen et al., 2013) to quantify the heritability of two distinct high-dimensional traits—stimulus-evoked BOLD time courses and functional connectivity profiles—across the cortex.

      Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      We believe that this question is getting at the difference between pairwise ISC (i.e., correlating one BOLD time course from one subject with that from another subject) and leave-one-subject-out ISC (i.e., correlating one BOLD time course from one subject with the corresponding average time course across all other subjects). We chose to use the pairwise ISC method because it allows us to capitalize on the information contained in the n<sup>2</sup> pairwise ISC matrix (whereas the other approach averages out meaningful information to yield a n<sup>1</sup> ISC matrix) and leverage a more sophisticated multidimensional heritability approach. Also, the leave-one-subject-out approach introduces additional issues re: handling family-level data (e.g., should we include a subject’s twin in the leave-one-subject-out average? If so, how should we handle subjects who don’t have a twin in the dataset, as averaging data from different numbers of subjects will lead to different ISC magnitudes? etc.).

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      We report p-values for heritability throughout our paper (e.g., stating that BOLD time courses are significantly heritable in 99% of parcels in Figure 2), and we believe that the reliability of our spatial maps across days of data collection (also quantified with p-values) further demonstrates the trustworthiness of our results. Finally, as we demonstrate in Figure S5, our sample size is more than sufficient to reliably detect small effects.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      In addition to driving shared neuronal responses (which are captured in BOLD signal oscillations <.1 Hz or so), movies also elicit shared cardiac, respiratory, and motion responses across participants at higher frequencies. Although we used a relatively conservative denoising approach here, we believe some of these non-neuronal signals are still present in our data; alternatively, it is also possible that these signals reflect “fast” BOLD responses at >.15 Hz (as discussed in 10.1016/j.neuroimage.2021.118658). In any case, the fact that information in this frequency band is considerably less heritable than information in slower frequency bands supports the idea that this band is noisier and suggests that our heritability results are driven by canonical neuronal activity-related BOLD signals.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      Although the decrease of 0.025 is small, we note that this constitutes around ~50% of BOLD time course heritability in some voxels (seen in comparison to Fig. 4C), and the spatial pattern of this result is quite consistent across days of data collection, indicating its reliability. Furthermore, the whole-brain distributions of results shown in Fig. 5B are clearly skewed towards negative values, indicating that controlling for NT partially reduces (or “explains”) BOLD time course heritability. Still, we agree that showing raw h<sup>2</sup> values in addition to the difference maps would be helpful for some readers and have added a corresponding supplementary figure (S12) which shows these.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      We did consider adding standard errors for these heritability estimates, but found that visualizing standard errors for each of the 153 unique network combinations in our heatmaps rendered the visualizations difficult to parse, and given that our hypotheses concerned global (e.g., hyperaligned vs. MSM-aligned) or network-level (e.g., sensory vs. associative) patterns, we focused on calculating standard errors/p-values for these analyses (although we note that dyad-level standard errors can be found in Fig. 6B, where they are clearly marginal compared to the group effects).

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      We agree that this result was relatively under-explored in our Discussion section and have added additional text (lines 851-855) to connect this result to recent work on arousal-dependent uniqueness of FC.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Do the authors have any ideas why we see this hotspot of heritability in pMTG/LOTC? It really jumps out in Figure 1A and Figure 2. The more posterior sensory MT+ area seems to drop when regressing out ISC in Figure 2B, but this pMTG area stays hot. Is there anything special about this kind of multimodal biological motion/action observation / social perception area (Pitcher & Ungerleider, 2021)? I don't think this is necessary to discuss in the manuscript, but I'm curious if the authors have any speculation.

      We are not certain as to why BOLD time courses in this parcel are particularly heritable- although this area is associated with biological motion, that particular function tends to be more right lateralized, and here we see nominally higher heritability in the left hemisphere. Per a Neurosynth review (and consistent with the left lateralization), we believe this may have more to do with speech processing, but a more definitive answer will require further investigation.

      (2) Page 3, line 127: "More information on these clips"-it might be worth saying a little bit more here just to make sure people understand that these are audiovisual clips, they include language, they're long enough to convey meaningful social and narrative information, etc.

      We agree and have added additional details on the clip composition to the relevant methods paragraph.

      (3) Figure 1 caption: can you add a sentence reminding readers what's going on with Day 1 and Day 2?

      We thank R1 for this suggestion and have added a sentence to this effect at this location.

      (4) Page 9, line 379: "although these more associative parcels do not encode a substantial amount of stimulus-specific information"-is this really true? I suspect these association areas still have decent ISCs, even if there are many processing stages downstream of the raw stimulus.

      Although these parcels are not the most synchronized by the stimulus, we agree that it is unfair (and vague) to say that they do not encode a substantial amount of stimulus-specific information. We have edited this sentence to make a more specific claim and highlight the relatively lower ISC in these parcels vs. more unimodal sensory areas.

      (5) Page 9, line 417: Can you unpack a bit more what you mean by "supra-BOLD frequency band"?

      Here, we refer to the fact that BOLD signals resulting from neuronal firing events have frequencies below ~.15 Hz (Josephs and Henson, 1999). We have added additional text and the Josephs and Henson citation to this line to further unpack this point.

      (6) Page 18, line 695: This discussion of how attention and gaze might partly shape response time series reminded me of recent work by Borovska & de Haas (2024)-might be worth citing.

      We are grateful to R1 for alerting us to this very relevant work and have included a reference to it in our discussion.

      (7) Page 19, line 755: I'm not sure I'd describe the hyperalignment results here as a "deleterious effects [on] heritability"-my reading was that hyperalignment allows you to say something more specific about heritability of function by allowing you to effectively factor out heritability effects that reduce to individual differences cortical topography; this seems like a good thing!

      We agree that “deleterious” was a poor word choice given its negative connotation, and have edited this sentence to read:

      “With this in mind, future studies investigating genetic correlations between brain function and behavioral variables may benefit from hyperalignment, as it can factor out individual-specific cortical topography and thus yield more precise estimates of functional heritability.”

      (8) I would love to see a ventral view in some of these plots! Not asking you to recreate the figures, but the ventral temporal cortex is an area of interest for many folks in the movie fMRI space (e.g., Haxby et al., 2011).

      We agree that ventral views would be of interest to some readers and have added the corresponding maps for our main results in supplementary figures S3 and S9.

      References:

      Borovska, P., & de Haas, B. (2024). Individual gaze shapes diverging neural representations. Proceedings of the National Academy of Sciences, 121(36), e2405602121. https://doi.org/10.1073/pnas.2405602121

      Haxby, J. V., Guntupalli, J. S., Connolly, A. C., Halchenko, Y. O., Conroy, B. R., Gobbini, M. I., Hanke, M., & Ramadge, P. J. (2011). A common, high-dimensional model of the representational space in human ventral temporal cortex. Neuron, 72(2), 404416. https://doi.org/10.1016/j.neuron.2011.08.026

      Pitcher, D., & Ungerleider, L. G. (2021). Evidence for a third visual pathway specialized for social perception. Trends in Cognitive Sciences, 25(2), 100-110. https://doi.org/10.1016/j.tics.2020.11.006

      Reviewer #2 (Recommendations for the authors):

      (1) To address the common core analytical problems listed under A), B), C), D), E), and basically throughout the methods:

      (a) Conduct permutations with exchangability restrictions to account for the pattern of dyad-relationships as e.g. implemented in PALM.

      (b) Control for age and sex covariates as covariates (e.g. as in SOLAR), rather than by matching.

      (c) Perform r-to-z transforms when conducting further analyses on correlations that assume normality.

      (d) For all analyses that assume normal distributions, e.g. in SOLAR and BrainSMASH, check that this is the case.

      We have explained how PALM is not suited for the study of effects that are defined at the dyad level (A), that we controlled for age and sex covariates in all our formal heritability analyses in our original submission (B), that we always performed r-to-z transforms when indicated in our original submission (C), and that our spatial permutation results don’t hinge on distributional differences (D).

      (2) Replace SEs derived from kacknife approach with those from SOLAR, or provide a comparison and motivation and/or demonstrate that SEs are correct.

      A more thorough explanation of the block jackknife procedure can be found in prior work introducing the multidimensional heritability method used here (Anderson et al., 2021).

      (3) Given problem (F & G):

      (a) Consider studying the parameters that drive the hyperalignment. They can be included as covariates in heritability analyses, and/or their heritability is of interest to understand the reasons for the heritability reduction post-hyperaligment.

      We agree that this would be interesting but the specific parameters that drive hyperalignment are beyond the scope of this study.

      (b) Include the alternative explanation of hyperalignment-induced noise in the discussion.

      We have added a figure showing that hyperalignment does not increase noise in ISC and explained here why “hyperalignment-induced noise” does not constitute a reasonable alternative explanation for our results.

      (4) Add heritability results for NT phenotypes.

      We have added heritability analyses for NT topography and (global) NT magnitude, as detailed above.

      (5) Motivate global signal removal, and acknowledge this process typically alters results substantially.

      We have added an explanation of our rationale for using GSR and shown in this response that it does not in fact substantially alter the results.

      (6) Rephrase and/or clarify the following:

      (a) "permutations quantify average differences" (under A).

      (b) "network combinations" and related analyses (under B & C).

      (c) why some analyses are separated per visit/day and others not (C).

      (d) methods and reasons for sample size estimation (C).

      We have rephrased or clarified all of the above.

      Reviewer #3 (Recommendations for the authors):

      (1) Participants should be recleared. I know HCP 7T data has 184 subjects. How can the authors have 176 twins and 690 unrelated subjects?

      As we reported in our Methods section, 178 subjects had complete movie-watching datasets, and 176 subjects had complete movie-watching and resting-state datasets. Of the 178 subjects with complete movie-watching data, we identified 690 age- and sex-matched dyads.

      (2) Figure 1. I don't find Figure S1A in Figure S1.

      We thank R3 for catching this error- we have amended this reference to read Fig. S1.

      (3) I could also suggest putting Figure 1 and Figure 2 together.

      We thank R3 for this suggestion- ultimately, we prefer to keep these figures separate to reinforce the difference between our dyadic similarity and formal heritability analyses.

    1. eLife Assessment

      This study presents important findings by identifying small molecules that can stabilize and refold missense-mutated VHL tumor suppressor protein, offering a potential therapeutic approach for clear cell renal cell carcinoma. The computational design approach is well-executed, but the evidence is incomplete due to insufficient demonstration that HIF2 downregulation occurs through on-target VHL rescue rather than off-target effects. Additional experiments with appropriate controls are needed to establish the specificity of the mechanism.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed some of comments raised in the previous round of review and have opted to proceed to a Version of Record without additional review.]

      Summary:

      This is an excellent and strong paper. The authors not only show the mechanisms of action of destabilizing mutations in VHL, but notably, they also go on to computationally design and experimentally test an inhibitor that restores wild-type pVHL function, offering starting points for a new class of kidney cancer drugs. The approach that the authors take here can be used to target destabilizing mutations in repressor proteins, common in diseases, including cancer.

      Strengths:

      This paper is the culmination of an extraordinary amount of work, over years, including method development and testing by a broad range of tools and experiments. It is thorough and comprehensive. It is also well-written and easy to follow.

    3. Reviewer #2 (Public review):

      Summary:

      Inactivating VHL mutations are common in clear cell renal cell carcinoma, and about half of those mutations unfold/destabilize the protein rather than directly interfering with critical protein-protein interactions. The authors identify a compound that can stabilize/refold mutant VHL and seemingly restore its ability to downregulate its major downstream targets.

      Strengths:

      The authors use a clever combination of virtual and cell-based screens, followed by suitable biophysical and cell-based validation assays, to arrive at a VHL refolder. This compound is suboptimal from an ADME point of view, but could be a starting point for further medicinal chemistry optimization. Success would have implications for other diseases linked to similar loss-of-function mutations.

      Weaknesses:

      In going from CP4 to CP4.29 the authors screened based on downregulation of HIF. This is logical but also introduces the danger of identifying chemicals that can downregulate HIF in an "off-target" manner i.e. non-specifically. It therefore essential to clearly show that CP4.29 downregulates steady-state levels of HIF and HIF target genes in cells with suitable (hydrophobic core) VHL mutants but not in isogenic cells lacking VHL.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We are most grateful to both reviewers for providing valuable feedback on our manuscript.

      Reviewer 1 had solely favorable comments, with no suggestions for revision.

      Reviewer 2 pointed out that experiment evaluating the effect of CP4 on pVHL half-life (originally included as Figure 3c) was difficult to evaluate because of CP4’s effect on pVHL abundance prior to cycloheximide treatment. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      Reviewer 2 also pointed out that experiment evaluating the effect of CP4.29 on HIF-2α half-life (originally included as Figure 4g) was not very compelling. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      We agree with Reviewer 2’s suggestion that additional experiments could further solidify that C4.29 downregulates HIF2 in a purely “on-target” manner, however we prefer to reserve such studies for the future.

      Reviewer 2 also made several valuable suggestions for the text itself (awkward wordings / citations / clearer figure legends). We appreciate this feedback and have updated the text accordingly.

    1. eLife Assessment

      This important study advances our understanding of the biomechanics of seed processing in birds by providing a comprehensive 3D kinematic analysis of coordinated bill and tongue movements across two species with contrasting biting forces. The evidence is convincing, combining high-speed XROMM with Bayesian statistical modeling in a rigorous and technically innovative framework that advances the understanding of avian feeding kinematics. Strengthening the statistical validation of qualitative claims, particularly for tongue-seed velocity relationships, and improving the accessibility of the probabilistic modeling framework would further solidify the conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      The authors quantified and compared the 3D kinematics of bill and tongue movements between two seed-eating bird species: one that specializes on soft seeds, and one that is more adapted to feeding on hard seeds. Their goal was to determine specifically what the role of the tongue was for processing (e.g., dehusking) seeds, and to understand how differences in biting strength between species affect other aspects of seed processing. The authors provided intricate (visual) details of seed processing movements, and showed how coordination between the tongue and cranial kinesis (i.e., mobility of the upper bill relative to the cranium) is both critically important for properly positioning seeds to enhance feeding efficiency. Many studies have detailed how seed-eating birds process seeds, but this study has elevated those to a new level of quantification and visualization for readers to fully experience firsthand. Furthermore, the authors established that the force-velocity trade-off that has been observed between bill functions (e.g., feeding and singing) is largely driven by the contractile properties of the muscles. The conclusions are well supported by the results, and the authors placed the results more broadly into the context of manual grasping, making the argument that these birds achieve high levels of dexterity with far fewer degrees of freedom, which could have potential biomimetic applications.

      Strengths:

      This study builds upon - and advances - our understanding of the feeding mechanics of seed-eating birds using cutting-edge 3-dimensional modeling and kinematics. Their quantitative analyses of upper and lower bill, tongue, and seed displacements are complemented by elegant visualizations of seed processing in each species. Their comprehensive Bayesian modeling statistical framework tackles the issue of small sample sizes (i.e., few subjects) with volumes of data for each (i.e., lots of sequential kinematic variables) that plague comparative biomechanics studies, principally because (a) it is difficult to gather these high resolution XROMM and muscle contractile data on more than just a few subjects, and (b) these data streams are inherently very large, as they are gathered at high frame and sampling rates. Furthermore, I believe their approach to statistically testing for differences between species sets a new standard for our field that could (perhaps should?) be implemented in other similar types of studies. Another strength is in how the results were packaged: each subsection indicated how the objectives were addressed, and there were concluding statements trailing each subsection that helped deliver the key takeaways.

      Weaknesses:

      A potential weakness is one that the authors themselves mentioned, regarding the body (and skull) size differences between species. Because gape size limits bite force, and given the force-velocity tradeoff in muscle function, there could be limitations on the rapid manipulation of relatively large seeds for similar reasons in the smaller finches. I see that the small finches appear to overcompensate in their beak rotations, but it's not clear how those compensatory movements might affect their seed processing kinematics with their preferred seed sizes. This does not nullify the authors' conclusions, but the results for the smaller finches might not be entirely representative of seed processing mechanics in smaller species.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates coordinated beak-tongue movements in seed manipulation, biting, and dehusking in songbirds. A comparative analysis of the seed-eating process in two songbird species with different biting forces, the domestic canary and Java sparrow, was conducted using high-speed XROMM with anatomical marker tracking and quantitative behavioral analysis. The authors have done a great job analyzing upper and lower beak rotation and translation, seed orientation and movement speed, and tongue kinematics.

      Strengths:

      The methodological approach of using high-speed (500 fps) X-ray reconstruction for 3D kinematic tracking in small animals is novel and powerful. It enables high temporal resolution tracking of orofacial movements and could potentially inspire future orofacial research in mammals, including mice and marmosets. Moreover, this study encompasses a wide range of anatomical components involved in seed manipulation behavior, including the upper and lower beak, the tongue, and jaw muscles. The behavioral quantification of these components is solid. The findings that both the upper and lower beaks contribute to seed processing, that the lower beak exhibits greater up-and-down and left-to-right flexibility than the upper beak during seed processing, and that the tongue plays an important role in transporting seeds into the mouth are all solid conclusions consistent with observations of bird feeding behavior. Nevertheless, it is valuable to confirm and quantitatively characterize these observations experimentally. The videos are excellent and very informative.

      Weaknesses:

      (1) The paper often resorts to qualitative descriptions (e.g., "a high positive correlation of tongue velocity and seed velocity", "Compared to positioning, the measured velocities of both seed and tongue were much lower") instead of providing exact quantitative measurements or statistical results. The authors stated that temporal autocorrelation biases standard statistical analyses (lines 205-210), but this rationale does not justify the absence of statistical validation. Suggestion: use appropriate methods for time-series data, such as a permutation test, to test the significance of correlations between variables and avoid false positives.

      (2) (Minor) The marker-tracking image shown in Figure 1B could benefit from the inclusion of a higher-contrast, zoomed-in frame of the head showing the metal markers without the red tracking points, alongside the same frame with the red tracking points overlaid, to provide readers with a clearer view of the X-ray image and the methodology and its precision.

      (3) (Minor: possibly soften the mechanistic claim). The proposed mechanism of lingual papillae on the tongue surface may aid food manipulation and food movement towards the posterior region of the mouth is interesting, yet the evidence describing their morphology is not strong enough to support the claim about their functional roles. Furthermore, the claim that papillae orientation affects food transport in lines 294-296 lacks supporting experimental evidence. In addition, the roles of extrinsic and intrinsic tongue muscles in controlling dexterous tongue shape changes and movements are not discussed.

    4. Author response:

      We would like to express our gratitude for the thorough evaluation of our manuscript by the editors and reviewers. We are grateful for the overall positive assessment. The suggestions for improvement are reasonable, and we are certain that addressing these points will improve the clarity, accessibility, and scientific integrity of the study. Thus, we plan to conduct a revision of the manuscript, addressing all the points raised. The most important planned adjustments are outlined below.

      (1) Improving the accessibility of the probabilistic modeling framework

      Reviewer 1 kindly stated that our Bayesian modeling framework for testing for species differences 'sets a new standard for our field.' As a new standard, however, the method should be explained in a more accessible way. Hence, we plan to provide additional explanations for the statistical workflow, e.g., by providing comprehensible visuals, to make the workflow easier to understand and easier to apply.

      (2) Statistical validation of qualitative claims

      We acknowledge that a statistical validation of qualitative claims regarding the relationship between seed and tongue movements and between upper and lower beak movements would considerably strengthen the validity of our findings. We thank Reviewer 2 for bringing permutation tests to our attention for quantifying the correlation between time series. Since permutation tests involving index-shuffling of one of the data sets are generally not valid for time-series data [1, 2], we'll consider a variant of a trial-swapping permutation test, such as a permute-match test [3]. Alternatively, the truncated time shift (TTS) test [2] might be an option, as also this method is valid for auto-correlated time series data. At this point, we can't tell yet which method we'll use for the revised manuscript. We need more time to assess the requirements of each method and evaluate which test is most appropriate to answer our specific research questions and best fits our kind of data.

      (3) Adjustments in the discussion

      Following the suggestion by Reviewer 1, we'll refine our discussion on the effects of skull size differences, putting more emphasis on the implications of potential effects for feeding kinematics in small species.

      Furthermore, as suggested by Reviewer 2, we'll soften our discussion on potential functions of lingual papillae in seed processing, as the current literature lacks experimental evidence for the claimed mechanistic roles.

      References

      (1) Yuan, A. E., & Shou, W. (2022). Data-driven causal analysis of observational biological time series. Elife, 11, e72518.

      (2) Yuan, A. E., & Shou, W. (2024). A rigorous and versatile statistical test for correlations between stationary time series. PLoS biology, 22(8), e3002758.

      (3) Yuan, A. E., & Shou, W. (2025). Permute-match tests: Detecting significant correlations between time series despite nonstationarity and limited replicates. eLife, 14.

    1. eLife Assessment

      This valuable study investigates the neural basis for recovery of complex wheel running behaviour following a unilateral spinal cord injury in mice. By combining behavioural analyses, whole-brain mapping, and tracing techniques, the authors provide incomplete evidence that new cortico-medullary connections can drive effective motor recovery. The paper could be strengthened with manipulations to establish causality, a more fine-grained analysis of the behaviour, and some reorganisation of how the data are presented and discussed.

    2. Reviewer #1 (Public review):

      Summary:

      The authors seek to understand and identify the neural plasticity that underlies recovery from precise unilateral hemi-pyramidotomy. The corticospinal tract is severed on one side in the pyramids below the exit of corticoreticular projections. Recovery from the injury is achieved with an intensive wheel running rehabilitation regime. The anatomical sites of plasticity, the importance of plasticity in different reticular areas<br /> to recovery, and the impact of the degree of plasticity observed on recovery as correlated predictors, are shown.

      Strengths:

      Refined anatomical analysis using mouse line and genetic and viral intersectional tracing identifies specific reticular targets of likely enhanced cortical control that correlate with recovery of locomotor skill.

      Weaknesses:

      (1) The study is correlational at this time. This does not undercut the value of the data and the identification of targets of plasticity achieved in the work.

      (2) Generalization of motor gains beyond locomotion was not tested. Reach-to-grasp tasks for feeding were not tested.

      (3) Some discussions and use of the terms fine motor and skilled motor are fuzzy, and the limitations of the study are not sufficiently clearly stated.

    3. Reviewer #2 (Public review):

      Summary:

      Bonanno and colleagues combine unilateral pyramidotomy, continuous voluntary complex-wheel running, whole-brain intersectional CSN tracing, and c-Fos mapping to ask whether rehabilitation reorganizes the supraspinal collaterals of the intact corticospinal tract neurons. The study is technically ambitious and competent, the uPyX + complex-wheel + intersectional-tracing + BrainJ combination is smart and interesting, the behavioral effect is convincing, and the blinding and exclusion criteria are explicit. The central anatomical finding - a CSN-specific, whole-brain projectome comparison with subregional LPGi/GiA/MdV granularity - is a legitimate contribution that builds on Asboth 2018. However, the strength of evidence does not support the strongest causal wording in the current abstract, significance statement, and parts of the discussion: the results remain correlational, the MdV-behavior correlation is modest, and its significance is sensitive to the unit of analysis. A major revision is recommended, primarily of framing and quantitative robustness, rather than because the central dataset is unconvincing.

      Strengths:

      (1) Technically ambitious and technically competent study addressing a relevant gap: brain-wide mapping of intact-CSN reorganization under continuous voluntary rehabilitation.

      (2) The combination of uPyX, complex-wheel running, intersectional tracing, and BrainJ whole-brain projection analysis is novel and well integrated.

      (3) Behavioral effect is convincing, blinding, and exclusion criteria are explicit.

      (4) The central anatomical finding (CSN-specific whole-brain projectome under rehab, with LPGi/GiA/MdV subregional resolution) is a legitimate contribution that builds on Asboth 2018. The closest recent works (Lemieux et al. 2024, Jeleva et al. 2026) study reticulospinal rather than CSN plasticity and are complementary rather than competing.

      Weaknesses:

      (1) Causal framing extends beyond what the current evidence supports.

      The abstract and significance statement present MdV as a potential mediator, or even a central locus, through which rehabilitation re-establishes descending control of the impaired limb. This is stronger than the evidence. What the paper shows is that CSN collateral projection density in MdV has a mild-to-medium correlation with behavioral recovery, and that this region is already known from prior work (Esposito 2014) to be relevant for skilled forelimb function. That is an interesting anatomical correlation, not a demonstration of mediation. No manipulation of MdV or of MdV-projecting CST terminals is performed; there is no silencing, no pathway-specific perturbation during rehabilitation, and no test showing that the identified sprouting is necessary for recovery. The limitations section acknowledges this, but the prominent claims do not.

      (2) The behavioral caveat on what is actually novel.

      The cleanest way to state what is genuinely new, clearer than the abstract itself, is this: when a CSN population loses part of its spinal target domain (via contralateral uPyX denervating the opposite cord), some CSNs from the opposite cortex appear to redirect growth into brainstem collaterals (LPGi, GiA, MdV). The compensation is plausibly sufficient to restore gross descending drive to the impaired forelimb, but most probably inadequate for the fractionated, cortico-motoneuronal fine-grain control that the direct CST normally provides. That distinction - recovery of drive and even skilled locomotor control vs. recovery of fine precision - is consistent with the ladder-rung improvements the paper reports (footfall counts are an integrated gross-placement metric) and with the skilled-reaching literature (Esposito 2014 and similar), which suggests precision grip and digit individuation would not be fully recovered by an MdV-centered detour. This note is also translationally important when we ask what humans consider fine motor control, which is mostly object manipulation. Relatedly, the ladder task is "skilled" in the operational sense that it requires cortical control, but the motor output measured (gross paw placement, overreach) is not fine motor function in the sense of digit individuation, grip force modulation, or pellet manipulation. "Skilled" here does not even mean *acquired* skill: classical skilled reaching in rodents involves explicit training to acquire a novel motor program, whereas here mice are only habituated. The brainstem-compensation hypothesis is more comfortable with restoring cortex-dependent gross placement than with restoring acquired fine-motor skills.

      (3) The anatomy sample is modest for the precision of the claims.

      Projection analysis rests on n = 9 pooled controls, n = 5 uPyX−Rehab, and n = 5 uPyX+Rehab. For a whole-brain subregion analysis, this is not a large dataset, even with the sensible restriction to the Wang et al. spinally-projecting set. The three medullary hits are plausible, but some of the most specific conclusions rely on a relatively small number of animals for its most specific claims. This matters especially for the MdV-behavior correlation.

      (4) Normalization enforces a zero-sum structure.

      Projection density is normalized to the total CST tract signal. This is a reasonable way to control for tracing variability, but it imposes a relative structure on the data: an apparent increase in one region may partly force an apparent decrease elsewhere. This may matter and has to be looked into by the authors, because the manuscript interprets decreased density in some other targets as meaningful redistribution.

      (5) The decision to merge PMn and MdV under a single "MdV" label needs more justification.

      Since the discussion relies on prior literature assigning skilled forelimb function to MdV proper, the reader needs to know whether the signal truly localizes there or whether it may partly reflect a neighboring region grouped under the same atlas label. Related to this, laterality would be very informative: since the proposed compensatory route is anatomically directional, showing whether the increased signal is preferentially located on the expected side of the medulla would strengthen the interpretation.

      (6) The c-Fos / Fig. 3 section goes beyond what the data directly support.

      The section "Complex-wheel running recruits intact corticospinal neurons" and the figure title "Rehabilitation functionally recruits intact CSNs" go beyond the actual observation, which is that a higher fraction of CSNs in M1 and M2 are c-Fos+ in runners than in non-runners. "Functionally" is not supported: c-Fos is a transcriptional marker of recent activity, not a functional readout; it does not show that the CSN's output is used to drive behavior. "Rehabilitation" is not supported either: the contrast is runners vs non-runners, applied uniformly across Sham and uPyX groups - healthy Sham+Rehab animals are on wheels for leisure, and the c-Fos effect is present in them too. The finding is difficult to interpret without thinking of the simpler framing ("moving mice have more motor cortex activity than resting mice"), with no control for generic arousal or ambulation. This section is the softest link in the causal chain running - CSN activity - medullary sprouting - recovery.

      (7) MdV-recovery correlation: unstated multiple-comparison correction and possible pseudoreplication.

      The correlation (R² ≈ 0.33, p ≈ 0.01) is the backbone of the paper's "causal" claim. Panels L/M/N test three correlations (LPGi, GiA, MdV vs forelimb footfall recovery); only MdV is reported as significant. The Figure 5 legend applies Tukey adjustment to the t-tests in A-C but makes no analogous statement for the correlations in L-N. A 3-test Bonferroni (α = 0.017) would not flip the MdV result, but disclosure is warranted, and the three tested regions were pre-selected from the significant group contrasts in A-C, which, to a statistician, would further shrink effective α. More importantly, the figure legend states that closed and open circles represent CFA- and RFA-traced values, respectively, which suggests the correlation treats the two tracer channels per mouse as independent datapoints - doubling the apparent n (≈ 20 from 10 uPyX mice), with the result of a higher significance than one would have at the mouse level.

      (8) Reporting issues.

      The reader would benefit from judging statistical choices such as those above directly from a data table rather than interpreting the authors' choices. The SciScore rightfully flags multiple missing components of transparent reporting: missing RRIDs, no code availability, limited data availability, and no power calculation, among others.

      Almost all these weaknesses can be addressed with a revision of the manuscript, especially in the framing of results.

      Conclusion:

      The core message - that rehabilitation is associated with a selective pattern of CSN collateral remodeling in the motor medulla, and that MdV projection density covaries with behavioral recovery - is defensible from the data and already a useful result. The current wording in parts of the abstract, significance statement, and discussion goes beyond this and implies a mechanistic conclusion (mediation, central locus, re-establishment of descending control) that the data do not yet establish. The manuscript would better match its evidence with "associated with", "correlates with", or "candidate locus" framing, unless a causal experiment is added.

    4. Reviewer #3 (Public review):

      Summary:

      In this study, Bonanno et al. show that after a lesion of the corticospinal tract (CST), rehabilitation running in a complex wheel drives improvement in skilled forelimb performance in mice. Mice with unilateral CST injury can perform gross motor tasks (locomotion) at the same level as the non-injured mice, but injured mice still have deficits in another task involving fine motor control. Thus, it is well-suited to test the efficacy of locomotion-based rehabilitation in fine motor control. Mice that voluntarily engaged in the rehabilitation protocol improved in the fine motor control task more than those mice that did not perform any rehabilitation. Highlighting the role of rehabilitation in the recovery of motor function after the lesion.

      The authors aimed to study rehabilitation-driven intact CST sprouting to supraspinal areas. They identified one area in the motor medulla where rehabilitation significantly changes the projection density from the intact cortical spinal neurons. Interestingly, this area has ipsilateral connections and thus could be a pathway to convey motor commands from the intact corticospinal tract to the denervated area. However, as the authors acknowledge in the discussion, they only found a correlation between the change in the synaptic projections from intact CST to the medulla and the recovery. Future work should study if indeed the area of the motor medulla identified here increases its ipsilateral projections to the denervated area, confirming the re-routing of motor commands from the intact cortico spinal tract to the denervated area. The paper is strong and, in general, claims are supported by the data.

      Strengths:

      In this study, Bonanno et al. show that after a unilateral corticospinal tract lesion (CST), locomotion rehabilitation can improve motor function and improvements generalized to tasks that require fine motor control. Moreover, it identifies a potential pathway that could be used for the intact corticospinal tract to convey motor commands to the denervated area. The pathway identified here could become a target for rehabilitation therapies.

      Weaknesses:

      As the authors acknowledge in the discussion of the study, the main limitation of this study is that the reorganization observed at the motor medulla is only correlational. Thus, it is possible that the adaptation to running with an injured limb of the intact CST to adapt to an injured limb rather than a re-routing of the intact CST inputs to the denervated area underlies the synaptic changes observed in the motor medulla.

      The statistical analysis could be better described.

      The generalization of skilled movement is limited to only locomotion tasks.

    1. eLife Assessment

      The worldwide decline in the health of coral reefs is well documented, and overgrowth by microbial consortia can be a contributing factor. Kelman and colleagues used metagenomic analysis to interrogate potential changes in phage-associated genes predicted to be involved in central carbon metabolism. The study addresses the hypothesis that metabolic genes associated with carbon metabolism that are encoded by viruses reflect the health of the corals. The study contributes a valuable perspective on the potential role of phages in coral health, although limitations of the data and analyses offer an exploratory examination rather than a definitive result. Overall, the evidence supporting the major findings is incomplete, in part because the conceptual model relies on qualitative assumptions rather than empirical data.

    2. Reviewer #1 (Public review):

      Summary:

      Microbialization (bacterial overgrowth) is a recognized component of degraded, eutrophied coral reefs where there is a shift from coral to algal dominance on the benthos. In addition, previous work has demonstrated that virus communities shift from a lytic strategy dominated (kill-the-winner) to a temperate (lysogenic) strategy dominated with reef microbialization. Kelman et al. sought to leverage previously published virus metagenomes produced from the water column of healthy and degraded coral reefs to assess virus community metabolic shifts. The authors also produce a conceptual model to demonstrate the potential impact of the observed metabolism shifts on reef fates.

      Strengths:

      The main strength of the manuscript is the findings from their metagenomic analyses and results. The virus metagenomes were produced using established approaches in the field and yield sufficient data per sample for their analyses. Interesting results regarding the shift in the types of genes from anaplerotic to cataplerotic provide the foundation for testable hypotheses to determine the magnitude of impact virus strategies have on reef health. The introduction is also well written and sets up the scene very well.

      Weaknesses:

      (1) The methods text currently omits important information related to the sampling design. It is not clear how many metagenomes are from healthy and degraded communities. This impacts the interpretability and robustness of the statistical results. Furthermore, it is unclear if analyses are based on assembled contigs or read-based alignments. Improving the clarity and organization of the Methods is essential for reproducibility.

      (2) Regarding the bioinformatics approach, normalization using the "percent known" approach within samples may not fully account for discovery bias related to sequencing depth. While Supplementary Table 1 shows variability in read counts, the lack of community-level metadata makes it difficult to determine if sequencing depth covaries with community type (healthy vs. degraded). The study would benefit from a rarefaction analysis or subsampling to ensure that gene frequency trends and Spearman correlations are biological signals rather than artifacts of sequencing effort.

      (3) The qualitative model in Figure 5 is positioned as evidence for the role of viruses in reef health, but it does not provide independent support for the authors' hypotheses. Since the model is parameterized using "arbitrary units" to reflect the authors' assumptions rather than being derived from the empirical metagenomic data, it serves as a helpful illustration of a hypothesis but not as a validation of the findings.

      (4) Results and discussion require revisions to improve readability and connectivity across sections. Ensuring a clear distinction between empirical data and model-based speculation would help the audience better appreciate the science.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Kelman and coauthors investigates how viral communities differ in the genes they encode in healthy and degraded coral reef ecosystems. Across 19 viral metagenomes from Central Pacific reefs, the authors assess the frequency of integration/excision genes as a proxy for viral community temperateness and ask whether genes associated with central carbon metabolism covary with signatures of temperateness. The main finding is that viral communities with more temperate-related genes encode more genes from the Entner-Doudoroff pathway and other reactions interpreted as anaplerotic, whereas more lytic-associated viral communities show greater representation of some pentose phosphate pathway, TCA, and redox-associated genes interpreted as cataplerotic. The authors propose a model based on these patterns in which lytic viral metabolism helps suppress bacterial overgrowth on healthy reefs, while temperate viral metabolism may promote microbialization on degraded reefs. The study addresses an interesting and potentially important concept - that viral auxiliary metabolic genes are important components of microbial communities and can affect ecosystem functioning. Linking viral metabolism to coral reef microbialization is a creative conceptual advance. The manuscript is clearly written, and the reported enrichment of anaplerotic genes in temperate-associated viromes is an interesting pattern that could motivate future work on how viral metabolic potential varies across reef states.

      Strengths:

      (1) The study connects viral lifestyle, central carbon metabolism, bacterial overgrowth, and reef degradation in a framework that could be useful for future studies of coral reef ecosystems and viral ecology. This is an interesting synthesis that links viral auxiliary metabolism to broader questions about microbialization and reef state.

      (2) The manuscript is generally clearly organized around a testable prediction: viral metabolic gene content should vary along a lytic-to-temperate viral community gradient. The reported enrichment of anaplerotic genes in viromes with a larger fraction of temperate viruses is a compelling result.

      (3) The authors highlight several virus-encoded metabolic genes that may not have been previously reported in viral datasets or genomes. If supported by further validation, these observations could expand the known repertoire of viral metabolic potential.

      (4) The modeling helps clarify the feedbacks the authors propose may connect viral lifestyle, bacterial metabolism, and coral reef degradation. It provides a foundation for generating hypotheses about how viral metabolic genes could influence reef microbial dynamics.

      Weaknesses:

      (1) The main limitation is that the evidence for several key claims remains indirect. The core analysis is based on correlations between metabolic gene frequencies and integration/excision-related genes. This does not demonstrate that the metabolic genes occur in temperate viral genomes, are physically linked to lysogeny genes, are expressed during infection, or alter host metabolism. Thus, the data support an association between VLP-associated metabolic annotations and a community-level temperateness proxy, but not a direct link between temperate phages and these metabolic functions.

      (2) It is important not to equate community-level gene frequencies with genome-level or infection-level metabolic programs. A virome may contain more anaplerotic genes overall, but that does not demonstrate that individual viruses reprogram their hosts in an anaplerotic manner nor that infection produces a net anaplerotic effect. Individual viruses may encode both anaplerotic and cataplerotic genes, and a smaller number of cataplerotic genes could have stronger metabolic consequences depending on expression, enzyme efficiency, pathway position, and host context. This is an important limitation that should be acknowledged and, if possible, addressed with contig- or genome-level analyses.

      (3) The ecological interpretation assumes that viral infection is strong enough to influence reef-scale bacterial population dynamics. However, the study does not directly measure infection frequency, lysis rates, viral production, burst size, lysogeny frequency, prophage induction, gene expression, or bacterial mortality. If viral mortality or lysogenic conversion were rare in these systems, the observed gene-frequency patterns could have limited ecosystem-level consequences. This makes claims about viral metabolism suppressing bacterial overgrowth, accelerating microbialization, or acting as a conservation lever more speculative than suggested.

      (4) There are statistical limitations related to the use of relative gene frequencies. Because genes are normalized as percentages of known genes, the data are compositional. Apparent increases in some categories may partly reflect decreases in others. Bootstrapped Spearman correlations are useful for assessing the robustness of these associations, but they do not address compositionality or multiple testing.

      (5) The anaplerotic/cataplerotic classification is central to the manuscript's conclusions and would benefit from more support. The framework is useful, but it depends on both annotation confidence and biochemical context. Sequence-similarity annotations alone may be vulnerable to misannotation, especially for central metabolic enzymes that share conserved domains across functionally distinct proteins. Stronger evidence that key genes contain key functional domains and/or are phylogenetically related to characterized enzymes would help support the proposed functions. In addition, many central carbon enzymes are reversible or context-dependent, so a clearer rationale for each classification would strengthen the interpretation.

      Overall, the manuscript presents a valuable hypothesis and highlights new ecological patterns in coral reef viral metagenomes, but falls short of the evidence needed for the strongest claims. The work would be strengthened by analyses that directly link metabolic genes to viral genomes or lysogeny markers, address compositional effects, validate key annotations, and more clearly distinguish observed gene-frequency associations from hypothesized effects on infection, host metabolism, and reef state.

    1. eLife Assessment

      This Review Article puts forth a normative theory for the grid cell representations found in the entorhinal cortex. It discusses a range of theoretical models and experimental findings, organizing them around a proposed framework in which grid cells are interpreted as biologically constrained, high-fidelity codes for path integration. This framing can be potentially interesting both for readers seeking a conceptual entry point into the grid cell literature and for those more generally interested in the promises and limitations of normative theories in neuroscience. Some logical gaps and points requiring conceptual or technical clarification were nonetheless identified. Moreover, the empirical support for the path-integration account is not yet as definitive as the manuscript's framing sometimes suggests. The review would thus be strengthened by clearer justification of key arguments and fuller discussion of biological complexities, model limitations, and competing interpretations. Some stylistic choices in how arguments and literature are sometimes rhetorically framed may lessen the review's appeal for key segments of its intended audience.